Method, apparatus, device and system for processing control messages in a collective communication system

By directly using query and notification messages to obtain the on-network computing capabilities of the switch network in the collective communication system, the problems of complex deployment and difficult maintenance of management processes in the existing technology are solved, and an easy-to-maintenance and efficient on-network computing solution is realized.

CN113835873BActive Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010760361.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-08
Filing Date
2020-07-31
Publication Date
2025-07-11
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

The management process of network computing solutions in existing integrated communication systems is complex and difficult to maintain, especially in large-scale networking, which affects system performance and maintenance efficiency.

Method used

By reusing the context of the collective communication system, querying the on-network computing capabilities of the switch network based on the control packet, avoiding duplicate resource creation and management process dependence, and achieving flexible and easy-to-maintenance on-network computing.

Benefits of technology

It simplifies the deployment and maintenance of on-network computing solutions, improves the ease of maintenance, flexibility and versatility of the system, and reduces costs, does not rely on remote direct data access network communication standards and additional configurations, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113835873B_ABST
    Figure CN113835873B_ABST
Patent Text Reader

Abstract

The present application provides a method for processing control messages in a collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes a first switch. The method includes: the first switch forwards a query message transmitted from a source node to a destination node. The query message is generated by the source node according to the context of the collective communication system. Then, the first switch forwards a notification message transmitted from the destination node to the source node. The notification message carries the in-network computing capabilities of the switch network. This method directly performs control message sending and receiving by reusing the context to query the in-network computing capabilities, so that subsequent service messages can perform INC offloading based on the in-network computing capabilities, avoiding the repeated creation and acquisition of related resources, decoupling the dependence on the control plane management process and the computing node daemon process, and providing an in-network computing solution with better maintainability, flexibility, and generality.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application submitted to the China National Intellectual Property Administration on June 8, 2020, with the application number 202010514291.1 and the application title "Method and Device for In-network Computing Control of Aggregate Communication", the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of information technology, and in particular, to a method, apparatus, device, system, and computer-readable storage medium for controlling message processing in an aggregate communication system. Background Art

[0003] With the continuous development of high performance computing (HPC) and artificial intelligence (AI) technologies, many new applications have emerged. Users are increasingly pursuing the ultimate in the execution efficiency and performance of various application scenarios. Among them, aggregate communication is the mainstream communication method for various application scenarios and also the future development trend. By replacing a large number of point-to-point operations with aggregate operations in aggregate communication, the running performance of applications can be improved.

[0004] In an aggregate communication system, when computing nodes execute aggregate operations, they often occupy a large amount of computing resources, such as a large amount of central processor unit (CPU) resources. Based on this, the industry has proposed an in-network computing (INC) solution. In-network computing specifically uses the extreme forwarding ability and strong computing ability of in-network computing devices such as switches to offload aggregate operations, thereby greatly improving the performance of aggregate operations and reducing the CPU load of computing nodes.

[0005] Currently, a typical in-network computing solution in the industry is to deploy an independent management process on a management node. This management process includes a subnet manager and an aggregation manager, and then obtain the network topology and in-network computing ability within the communication domain through the above management process, and perform INC offloading on subsequent service messages based on the network topology and in-network computing ability.

[0006] However, the deployment process of the management process is very complex and the maintenance difficulty is relatively high. In large-scale network deployments, the deployment complexity and maintenance difficulty are even more prominent. Based on this, the industry urgently needs to provide a more concise and efficient in-network computing solution to optimize the performance of the aggregate communication system. Summary of the Invention

[0007] The present application provides a method for processing control messages in a collective communication system. This method directly performs the sending and receiving of control messages such as query messages and notification messages by reusing the context of the collective communication system, and queries the in-network computing capabilities based on the control messages, so that subsequent service messages can perform INC offloading based on the in-network computing capabilities, avoiding the repeated creation and acquisition of relevant resources and decoupling the dependence on the control plane management process and the computing node daemon process. The in-network computing solution implemented by this method is more maintainable, flexible, and general. The present application also provides a device, a device, a system, a computer-readable storage medium, and a computer program product corresponding to the above method.

[0008] In a first aspect, the present application provides a method for processing control messages in a collective communication system. The collective communication system includes a switch network (also referred to as a switch fabric) and multiple computing nodes. Among them, the switch network refers to a network formed by at least one switch. The switch network may include one switch or multiple switches. The switch network can also be divided into a single-layer switch network and a multi-layer switch network according to the network architecture.

[0009] The single-layer switch network includes a single-layer switch, that is, an access layer switch. The single-layer switch includes one or more switches. Each switch in the single-layer switch can be directly connected to a computing node, thereby connecting the computing node to the network.

[0010] The multi-layer switch network includes an upper-layer switch and a lower-layer switch. Among them, the upper-layer switch refers to a switch connected to the switch, and the upper-layer switch usually is not directly connected to the computing node. The lower-layer switch refers to a switch that can be directly connected to the computing node, and thus is also referred to as an access layer switch. For example, the multi-layer switch network can be a leaf spine architecture. The upper-layer switch is a spine switch, and the lower-layer switch is a leaf switch. The spine switch is no longer the large chassis switch in the three-layer architecture, but a high-port-density switch. The leaf switch, as the access layer, can provide network connections to computing nodes such as terminals and servers, and at the same time connect to the spine switch upstream.

[0011] The switch network includes a first switch. The first switch can be a switch in the above single-layer switch network or a switch in the multi-layer switch network, such as a leaf switch or a spine switch, etc. An application that supports collective communication can reuse the context of the collective communication system to initiate a control message process to query information such as the in-network computing capabilities of the switch network, so as to provide assistance for in-network computing (computing offloading).

[0012] Specifically, the source node can generate a query message according to the context of the collective communication system. The query message is used to request a query of the online computing capabilities of the switch network. The first switch can forward the query message transmitted from the source node to the destination node. The source node and the destination node are specifically different computing nodes in the collective communication system. The destination node can generate a notification message, and the notification message carries the online computing capabilities of the switch. The first switch can forward the notification message transmitted from the destination node to the source node to notify the source node of the online computing capabilities of the switch network.

[0013] Among them, the online computing capabilities of the switch network include the online computing capabilities of one or more switches through which the query message passes. One or more computing nodes of the collective communication system can act as source nodes and send query messages to the destination node. When these query messages pass through all the switches of the switch network, the online computing capabilities of the switch network returned by the destination node through the notification message refer to the online computing capabilities of all the switches of the switch network. When these query messages pass through some of the switches of the switch network, the online computing capabilities of the switch network returned by the destination node through the notification message specifically refer to the online computing capabilities of the above-mentioned part of the switches.

[0014] This method multiplexes the context of the collective communication system to directly send and receive control messages such as query messages and notification messages, and queries the online computing capabilities based on the control messages, so that subsequent service messages can perform 1NC offloading based on the online computing capabilities, avoiding the repeated creation and acquisition of relevant resources, and decoupling the dependence on the control plane management process and the computing node daemon process. Based on this, the online computing solution provided by the embodiments of this application is more maintainable, flexible and general.

[0015] Furthermore, this method supports multiplexing existing network protocol channels, such as Ethernet channels, without relying on communication standards, without adopting remote direct memory access (RDMA) network communication standards (infiniband, IB), and without additional configuration of IB switches, thus greatly reducing the cost of the online computing solution.

[0016] Moreover, this method also does not require running a daemon process on the computing node. Only an INC dynamic library (INC lib) needs to be provided, and the specified application programming interface (API) in the INC lib is called within the collective operation communication domain to implement the control message service logic.

[0017] In some possible implementation manners, when a query message passes through a first switch, the first switch may add its in-network computing capability to the query message. For example, the in-network computing capability of the switch may be added to the query field of the query message. Then, the first switch forwards the query message with the in-network computing capability of the first switch added to a destination node. Correspondingly, the destination node aggregates the in-network computing capabilities of the first switch based on the query message with the in-network computing capability of the first switch added, obtains the in-network computing capability of the switch network, and then carries the in-network computing capability of the switch network in a notification message.

[0018] In this way, the in-network computing capability of the switch network is achieved through a simple and efficient method, which helps the in-network computing solution for collective communication.

[0019] In some possible implementation manners, the in-network computing capability of the first switch includes the collective operation types supported by the first switch and / or the data types supported by the first switch. The first switch adds the collective operation types supported by the first switch and / or the data types supported by the first switch to the query message, so that a computing node can determine whether to perform computing offloading on the first switch according to the collective operation types and data types supported by the switch, thereby realizing in-network computing.

[0020] Among them, the collective operation types may include any one or more of the following: broadcast from one member to all members in a group, collection of data from all members by one member, scattering of data from one member to all members in a group, scatter / gather data operation among all members in a group, global reduction operation, combined reduction and scattering operations, and search operation on all members in a group. Herein, a member refers to a process in a process group.

[0021] The data types may include any one or more of byte, 16-bit integer (short), 32-bit integer (int), 64-bit integer (long), floating point type (float), double-precision floating point type (double), boolean type (boolean), and character type (char), etc.

[0022] In this way, the computing node can determine the in-network computing policy according to the set operation types supported by the first switch and / or the data types supported by the first switch. Specifically, the computing node can compare the set operation type of the current collective communication with the set operation types supported by the first switch, and compare the data type of the current collective communication with the data types supported by the first switch. When the set operation types supported by the first switch include the set operation type of the current collective communication, and the data types supported by the first switch include the data type of the current collective communication, the computing node can offload the computation to the first switch; otherwise, the computing node does not offload the computation to the first switch. This can avoid additional operations by the computing node when the first switch does not support the set operation type or data type of the current collective communication, improving the efficiency of the computing node.

[0023] In some possible implementation manners, the in-network computing capability of the first switch includes the size of the remaining available in-network computing resources of the first switch. Among them, the size of the remaining available in-network computing resources of the first switch can be characterized by the maximum number of concurrent hosts of the first switch, that is, the local group size.

[0024] Among them, the in-network computing capability of the first switch can include any one or more of the set operation types supported by the first switch, the data types supported by the first switch, and the size of the remaining available in-network computing resources of the first switch. When the first switch default supports various set operation types, the in-network computing capability of the first switch may not include the set operation types supported by the first switch. When the first switch default supports various data types, the in-network computing capability of the first switch may not include the data types supported by the first switch.

[0025] The first switch adds the size of the remaining available in-network computing resources of the first switch to the query message, so that the computing node can determine the in-network computing policy according to the size of the remaining available in-network computing resources. For example, offloading all the computations to the first switch, or offloading part of the computations to the first switch, etc., so as to make full use of the in-network computing resources of the first switch.

[0026] In some possible implementation manners, the first switch may also establish an entry according to the hop count of the query message, and this entry is used for the switch to perform calculation offloading on service messages. Specifically, in the service message process, the first switch can identify whether the service message is an in-network calculation message. If so, it matches the in-network calculation message with the entry established in the control message process. If the match is successful, it performs calculation offloading on the service message. Thus, real-time allocation of in-network calculation resources can be achieved. Further, after the collective communication is completed, the above-mentioned entry can be cleared to release the in-network calculation resources. In this way, the resource utilization rate can be optimized.

[0027] In some possible implementation manners, the first switch is directly connected to the source node and the destination node. Correspondingly, the first switch can receive the query message sent by the source node, forward the query message to the destination node, and then the first switch receives the notification message sent by the destination node and forwards the notification message to the source node. Since the topology of the switch network is relatively simple, the sending and receiving of the query message and the notification message can be achieved through one forwarding, which improves the efficiency of obtaining the in-network calculation ability.

[0028] In some possible implementation manners, the switch network further includes a second switch and / or a third switch. Among them, the second switch is used to connect the first switch and the source node, and the third switch is used to connect the first switch and the destination node. The second switch can be one switch or multiple switches. Similarly, the third switch can also be one switch or multiple switches.

[0029] When the switch network includes a second switch and does not include a third switch, the first switch receives the query message forwarded by the second switch, forwards the query message to the destination node, and then the first switch receives the notification message sent by the destination node and forwards the notification message to the second switch.

[0030] When the switch network includes a third switch and does not include a second switch, the first switch receives the query message sent by the source node, forwards the query message to the third switch, and then the first switch receives the notification message forwarded by the third switch and forwards the notification message to the source node.

[0031] When the switch network includes both a second switch and a third switch, the first switch receives the query message forwarded by the second switch, forwards the query message to the third switch, and then the first switch receives the notification message forwarded by the third switch and forwards the notification message to the second switch.

[0032] The first switch can forward packets to other switches, and then forward the packets through other switches, so as to indirectly transmit the query packet from the source node to the destination node and transmit the notification packet from the destination node to the source node, thereby obtaining the in-network computing ability of the switch network by sending and receiving control packets.

[0033] Among them, the first switch and the second switch can be vertically connected, that is, the first switch and the second switch are switches of different levels. The first switch and the second switch can also be horizontally connected, that is, the first switch and the second switch can be switches of the same level, such as switches in the access layer. Similarly, the first switch and the third switch can also be vertically connected or horizontally connected.

[0034] In some possible implementation manners, the switch network includes a single-layer switch, such as a top-of-rack (ToR) switch. Therefore, the first switch is the above-mentioned single-layer switch. Thus, the interconnection between computing nodes such as servers and the first switch in the cabinet can be realized. When the first switch notifies a packet, it directly forwards the notification packet to the source node, having relatively high communication performance.

[0035] In some possible implementation manners, the switch network includes an upper-layer switch and a lower-layer switch. For example, the switch network can be a leaf-spine architecture, including an upper-layer switch, that is, a spine switch, located in the upper layer, and a lower-layer switch, that is, a leaf switch, located in the lower layer. The first switch can be one of the lower-layer switches, such as a leaf switch.

[0036] Specifically, the first switch can determine a target switch from the upper-layer switches according to the size of the remaining available in-network computing resources of the upper-layer switch, then add the size of the remaining available in-network computing resources of the target switch to the notification packet, and then the first switch forwards the notification packet added with the size of the remaining available in-network computing resources of the target switch to the source node. In this way, in the subsequent service packet process, the computing node can also determine the in-network computing policy based on the size of the remaining available in-network computing resources of the target switch, specifically, the policy of computing offloading on the target switch.

[0037] In some possible implementation manners, the first switch can determine a target switch from the upper-layer switches according to the size of the remaining available in-network computing resources of the upper-layer switch by using a load balancing policy. In this way, it can avoid the upper-layer switch from being overloaded and affecting the collective communication performance.

[0038] In some possible implementation manners, when the first switch is a lower-layer switch, it may also send a switch query message to the upper-layer switch. The switch query message is used to query the size of the remaining available resources for in-network computing of the upper-layer switch, and then receive a switch notification message sent by the upper-layer switch. The switch notification message is used to notify the size of the remaining available resources for in-network computing of the upper-layer switch. In this way, it can provide a reference for the lower-layer switch to determine the target switch.

[0039] In some possible implementation manners, when the first switch is an upper-layer switch, it may also receive a switch query message sent by the lower-layer switch. The switch query message is used to query the size of the remaining available resources for in-network computing of the first switch, and then send a switch notification message to the lower-layer switch. The switch notification message is used to notify the size of the remaining available resources for in-network computing of the first switch. In this way, by sending and receiving the switch query message and the switch notification message, the size of the remaining available resources for in-network computing of the upper-layer switch is obtained, providing a reference for the lower-layer switch to determine the target switch.

[0040] In some possible implementation manners, the context of the collective communication system includes the context of the application or the context of the communication domain. By reusing these contexts, the repeated creation and acquisition of related resources can be avoided, and the dependence on the control plane management process and the computing node daemon process is decoupled.

[0041] In some possible implementation manners, the multiple computing nodes include a primary node and at least one sub-node. Among them, the source node may be the sub-node, and correspondingly, the destination node is the primary node. In some embodiments, the source node may also be the primary node, and the destination node may also be the sub-node.

[0042] In a second aspect, the present application provides a method for processing control messages in a collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch, and the multiple computing nodes include a first computing node and a second computing node.

[0043] Specifically, the first computing node receives a query message forwarded by one or more switches in the switch network. The query message is used to request to query the in-network computing ability of the switch network. The query message is generated by the second computing node according to the context of the collective communication system. Then, the first computing node generates a notification message according to the query message. The notification message carries the in-network computing ability of the switch network. Next, the first computing node sends the notification message to the second computing node.

[0044] This method reuses the context of the collective communication system to directly send and receive control messages such as query messages and notification messages, and queries the in-network computing capabilities based on the control messages, so that subsequent service messages can perform INC offloading based on the in-network computing capabilities, avoiding the repeated creation and acquisition of relevant resources and decoupling the dependence on the control plane management process and the computing node daemon process. Based on this, the in-network computing solution provided by the embodiments of this application is more maintainable, flexible, and general.

[0045] In some possible implementation manners, the in-network computing capabilities of the switch network carried in the notification message are specifically obtained from the query messages forwarded by the one or more switches. Each time a query message passes through a switch, the switch adds its own in-network computing capabilities to the query message. In this way, the computing node can obtain the in-network computing capabilities of the switch network by sending query messages and receiving notification messages, avoiding the repeated creation and acquisition of relevant resources and decoupling the dependence on the control plane management process and the computing node daemon process.

[0046] In some possible implementation manners, the query messages forwarded by the switch include the in-network computing capabilities of the switch added by the switch. The first computing node can obtain the in-network computing capabilities of the switch network based on the in-network computing capabilities of the one or more switches in the query messages forwarded by the one or more switches. In this way, the in-network computing capabilities of the switch network can be obtained through simple sending, receiving, and processing of control messages.

[0047] In some possible implementation manners, the first computing node is the master node or the slave node. When the first computing node is the master node, the second computing node can be the slave node. When the first computing node is the slave node, the second computing node can be the master node.

[0048] In a third aspect, this application provides a method for processing control messages in a collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch, and the multiple computing nodes include a first computing node and a second computing node.

[0049] Specifically, the second computing node generates a query message according to the context of the collective communication system. The query message is used to request to query the in-network computing capabilities of the switch network. Then, the second computing node sends the query message to the first computing node through one or more switches in the switch network. Next, the second computing node receives a notification message forwarded by the first computing node through the one or more switches. The notification message is generated by the first computing node according to the query message and carries the in-network computing capabilities of the switch network.

[0050] This method reuses the context of the collective communication system to directly send and receive control messages such as query messages and notification messages, and queries the in-network computing capabilities based on the control messages, so that subsequent service messages can perform INC offloading based on the in-network computing capabilities, avoiding the repeated creation and acquisition of relevant resources, and decoupling the dependence on the control plane management process and the computing node daemon process. Based on this, the in-network computing solution provided by the embodiments of the present application is more maintainable, flexible and general.

[0051] In some possible implementation manners, the in-network computing capabilities of the switch network are obtained by the first computing node according to the in-network computing capabilities of the one or more switches in the query message forwarded by the one or more switches. In this way, it can help the computing node obtain the in-network computing capabilities of the switch network by sending and receiving control messages.

[0052] In some possible implementation manners, the second computing node is the master node or the slave node. When the second computing node is the master node, the first computing node can be the slave node. When the second computing node is the slave node, the first computing node can be the master node.

[0053] Fourthly, the present application provides a control message processing device in a collective communication system. The collective communication system includes a switch network and multiple computing nodes, the switch network includes a first switch, and the device includes:

[0054] A communication module, configured to forward a query message transmitted from a source node to a destination node, where the query message is used to request to query the in-network computing capabilities of the switch network, the query message is generated by the source node according to the context of the collective communication system, and the source node and the destination node are different nodes among the multiple computing nodes;

[0055] The communication module is further configured to forward a notification message transmitted from the destination node to the source node, where the notification message carries the in-network computing capabilities of the switch network.

[0056] In some possible implementation manners, the device further includes:

[0057] A processing module, configured to add the in-network computing capabilities of the first switch to the query message when receiving the query message;

[0058] The communication module is specifically configured to:

[0059] Forward the query message added with the in-network computing capabilities of the first switch to the destination node.

[0060] In some possible implementation manners, the in-network computing capability of the first switch includes the set operation types and / or data types supported by the first switch.

[0061] In some possible implementation manners, the in-network computing capability of the first switch includes the size of the remaining available resources for in-network computing of the first switch.

[0062] In some possible implementation manners, the apparatus further includes:

[0063] A processing module, configured to establish an entry according to the hop count of the query message, where the entry is used for the first switch to perform computing offloading on service messages.

[0064] In some possible implementation manners, the first switch is directly connected to the source node and the destination node;

[0065] The communication module is specifically configured to:

[0066] Receive a query message sent by the source node, and forward the query message to the destination node;

[0067] Receive a notification message sent by the destination node, and forward the notification message to the source node.

[0068] In some possible implementation manners, the switch network further includes a second switch and / or a third switch, where the second switch is used to connect the first switch and the source node, and the third switch is used to connect the first switch and the destination node;

[0069] The communication module is specifically configured to:

[0070] Receive a query message sent by the source node, and forward the query message to the third switch;

[0071] Receive a notification message forwarded by the third switch, and forward the notification message to the source node; or,

[0072] Receive a query message forwarded by the second switch, and forward the query message to the destination node;

[0073] Receive a notification message sent by the destination node, and forward the notification message to the second switch; or,

[0074] Receive a query message forwarded by the second switch, and forward the query message to the third switch;

[0075] Receive a notification message forwarded by the third switch, and forward the notification message to the second switch.

[0076] In some possible implementations, the switch network includes a single-layer switch, and the first switch is the single-layer switch;

[0077] The communication module is specifically configured to:

[0078] Forward the notification message to the source node.

[0079] In some possible implementations, the switch network includes an upper-layer switch and a lower-layer switch, and the first switch is the lower-layer switch;

[0080] The apparatus further includes:

[0081] A processing module, configured to determine a target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing of the upper-layer switch, and add the size of the remaining available resources for in-network computing of the target switch to the notification message;

[0082] The communication module is specifically configured to:

[0083] Forward the notification message added with the size of the remaining available resources for in-network computing of the target switch to the source node.

[0084] In some possible implementations, the processing module is specifically configured to:

[0085] Determine a target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing by using a load balancing strategy.

[0086] In some possible implementations, the communication module is further configured to:

[0087] Send a switch query message to the upper-layer switch, where the switch query message is used to query the size of the remaining available resources for in-network computing of the upper-layer switch;

[0088] Receive a switch notification message sent by the upper-layer switch, where the switch notification message is used to notify the size of the remaining available resources for in-network computing of the upper-layer switch.

[0089] In some possible implementations, the switch network includes an upper-layer switch and a lower-layer switch, and the first switch is the upper-layer switch;

[0090] The communication module is further configured to:

[0091] Receive a switch query message sent by the lower-layer switch, where the switch query message is used to query the size of the remaining available resources for in-network computing of the first switch;

[0092] Send a switch notification message to the lower-layer switch, where the switch notification message is used to notify the lower-layer switch of the size of the remaining available resources for in-network computing of the first switch.

[0093] In some possible implementation manners, the context of the collective communication system includes the context of an application or the context of a communication domain.

[0094] In some possible implementation manners, the multiple computing nodes include a primary node and at least one sub-node;

[0095] The source node is the sub-node, and the destination node is the primary node; or,

[0096] The source node is the primary node, and the destination node is the sub-node.

[0097] In a fifth aspect, the present application provides a control message processing device in a collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch, and the multiple computing nodes include a first computing node and a second computing node. The device includes:

[0098] A communication module, configured to receive a query message forwarded by one or more switches in the switch network, where the query message is used to request a query of the in-network computing ability of the switch network, and the query message is generated by the second computing node according to the context of the collective communication system;

[0099] A generation module, configured to generate a notification message according to the query message, where the notification message carries the in-network computing ability of the switch network;

[0100] The communication module is further configured to send the notification message to the second computing node.

[0101] In some possible implementation manners, the query message forwarded by the switch includes the in-network computing ability of the switch added by the switch;

[0102] The generation module is specifically configured to:

[0103] Obtain the in-network computing ability of the switch network according to the in-network computing ability of the one or more switches in the query message forwarded by the one or more switches;

[0104] Generate a notification message according to the in-network computing ability of the switch network.

[0105] In some possible implementation manners, the device is deployed on the first computing node, and the first computing node is a primary node or a sub-node.

[0106] Sixth aspect, the present application provides a control message processing device in a collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The multiple computing nodes include a first computing node and a second computing node. The device includes:

[0107] A generation module, configured to generate a query message according to the context of the collective communication system, where the query message is used to request a query of the online computing capabilities of the switch network;

[0108] A communication module, configured to send the query message to the first computing node through one or more switches in the switch network;

[0109] The communication module is further configured to receive a notification message forwarded by the first computing node through the one or more switches. The notification message carries the online computing capabilities of the switch network, and the notification message is generated by the first computing node according to the query message.

[0110] In some possible implementation manners, the device is deployed on the second computing node, and the second computing node is a master node or a slave node.

[0111] Seventh aspect, the present application provides a switch. The switch includes a processor and a memory.

[0112] The processor is configured to execute instructions stored in the memory, so that the switch executes the method described in the first aspect or any implementation manner of the first aspect of the present application.

[0113] Eighth aspect, the present application provides a computing node. The computing node includes a processor and a memory;

[0114] The processor is configured to execute instructions stored in the memory, so that the computing node executes the method described in the second aspect or any implementation manner of the second aspect of the present application.

[0115] Ninth aspect, the present application provides a computing node. The computing node includes a processor and a memory;

[0116] The processor is configured to execute instructions stored in the memory, so that the computing node executes the method described in the third aspect or any implementation manner of the third aspect of the present application.

[0117] Tenth aspect, the present application provides a collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes a first switch. The multiple computing nodes include a first computing node and a second computing node.

[0118] The second computing node is configured to generate a query message according to the context of the collective communication system, where the query message is used to request a query of the on-network computing capabilities of the switch network;

[0119] The first switch is configured to forward the query message transmitted by the second computing node to the first computing node;

[0120] The first computing node is configured to generate a notification message according to the query message, where the notification message carries the on-network computing capabilities of the switch network;

[0121] The first switch is further configured to forward the notification message transmitted by the first computing node to the second computing node.

[0122] In an eleventh aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions direct a device to execute the control message processing method in the collective communication system according to any one of the implementation manners in the first aspect, the second aspect, or the third aspect.

[0123] In a twelfth aspect, the present application provides a computer program product including instructions, which, when running on a device, cause the device to execute the control message processing method in the collective communication system according to any one of the implementation manners in the first aspect or in the first aspect, the second aspect, or the third aspect.

[0124] Based on the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0125] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below.

[0126] Figure 1 It is an architecture diagram of a collective communication system provided by an embodiment of the present application;

[0127] Figure 2 It is a schematic diagram of on-network computing in a collective communication system provided by an embodiment of the present application;

[0128] Figure 3 It is a schematic structural diagram of a switch in a collective communication system provided by an embodiment of the present application;

[0129] Figure 4 It is a schematic structural diagram of a computing node in a collective communication system provided by an embodiment of the present application;

[0130] Figure 5Flowchart of a method for processing control messages in a collective communication system provided by an embodiment of the present application;

[0131] Figure 6 Schematic structural diagram of a query message in a collective communication system provided by an embodiment of the present application;

[0132] Figure 7 Schematic structural diagram of a notification message in a collective communication system provided by an embodiment of the present application;

[0133] Figure 8 Interaction flowchart of a method for processing control messages in a collective communication system provided by an embodiment of the present application;

[0134] Figure 9 Interaction flowchart of a method for processing control messages in a collective communication system provided by an embodiment of the present application;

[0135] Figure 10 Schematic structural diagram of a control message processing device in a collective communication system provided by an embodiment of the present application;

[0136] Figure 11 Schematic structural diagram of a control message processing device in a collective communication system provided by an embodiment of the present application;

[0137] Figure 12 Schematic structural diagram of a control message processing device in a collective communication system provided by an embodiment of the present application;

[0138] Figure 13 Schematic structural diagram of a collective communication system provided by an embodiment of the present application. Detailed implementation manners

[0139] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0140] First, some technical terms involved in the embodiments of the present application are introduced.

[0141] High performance computing (HPC) refers to using the aggregated computing power of a large number of processing units for computing, so as to solve complex problems, such as weather prediction, oil exploration, nuclear explosion simulation, etc. Among them, the aggregated computing power of a large number of processing units can be the aggregated computing power of multiple processors in a single machine, or the aggregated computing power of multiple computers in a cluster.

[0142] Artificial intelligence (AI) refers to computer programs running on a computer, enabling the computer to have the effects of human intelligence, thereby assisting or replacing humans in solving problems. For example, artificial intelligence can be used to achieve automated image detection, image recognition, audio detection, video surveillance, etc.

[0143] Artificial intelligence generally includes two implementation methods. One implementation method is the engineering approach, that is, using traditional programming techniques to make the computer present intelligent effects without considering whether it is the same as the methods used by humans or animal organisms. One implementation method is the modeling approach, specifically using the same or similar methods as those used by humans or biological organisms to make the computer present intelligent effects.

[0144] In some examples, the modeling approach may include simulating the genetic - evolution mechanism of humans or organisms based on the genetic algorithm (GA), or may also include simulating the activity patterns of nerve cells in the human or biological brain based on the artificial neural network (ANN).

[0145] With the continuous development of HPC and AI, some new applications have emerged. Users are increasingly pursuing the ultimate in the execution efficiency and performance of these applications. Based on this, the industry has introduced collective communication, replacing a large number of point - to - point operations with collective operations in collective communication, thereby improving the performance of applications.

[0146] Collective communication refers to organizing a communicator to serve a group of communicating processes and completing specific communication operations among the processes. Among them, a group of communicating processes form a process group, and the communicator (which can also be called a communication sub - entity) comprehensively describes the relationship among the communicating processes. The communicator specifically includes: process group, context, and topology, etc. Among them, the context refers to the environment when the process is executed. The topology refers to the distribution of the computing nodes where the process is executed.

[0147] The environment when the process is executed specifically refers to various variables and data on which the process depends, including register variables, files opened by the process, and memory information, etc. The context is essentially a snapshot of the environment, and this snapshot is an object used to save the state. Most of the functions written in a program are not independently complete. When using a function to complete the corresponding function, it is very likely to require the support of other external environment variables. The context is to assign values to the variables of the external environment so that the function can run correctly.

[0148] Each process is objectively unique and usually has a unique process identifier (process id, pid). The same process can belong to only one process group or multiple process groups (the process has its own number in different process groups, i.e., rank number). When the same process belongs to multiple process groups, considering the one-to-one correspondence between the process group and the communication domain, this process can also belong to different communication domains.

[0149] The specific operations (i.e., collective operations) between processes in a collective communication system are mainly for data distribution and synchronization operations. Collective communication based on the Message Passing Interface (MPI) generally includes two communication modes, one is a one-to-many communication mode, and the other is a many-to-many communication mode.

[0150] Among them, the communication operations in the one-to-many mode can include broadcast from one member to all members in the group, gather data from all members by one member, scatter data from one member to all members in the group, etc. The communication operations in the many-to-many mode can include scatter / gather data operations between all members in the group, global reduction operations, combined reduction and scatter operations, search operations on all members in the group, etc. Among them, reduction means dividing a batch of data into a smaller batch of data through a function. For example, reducing the elements of an array to a single number through an addition function.

[0151] An application can perform collective communication in multiple different communication domains. For example, the application is distributedly deployed on computing nodes 1 to N. The process groups on computing nodes 1 to K can form a communication domain, and the process groups on computing nodes K + 1 to N can form another communication domain, where N is a positive integer greater than 3 and K is a positive integer greater than or equal to 2. These two communication domains have their own contexts. Similarly, the application itself also has a context. The context of the application refers to the environment when the application is executed. The context of the application can be regarded as the global context, and the context of the communication domain can be regarded as the local context within the communication domain.

[0152] In-network computing (INC) is a key optimization technology for collective communication proposed in the industry. Specifically, in-network computing means using the extreme forwarding ability and strong computing ability of switches to offload collective operations, thereby greatly improving the performance of collective operations and reducing the load on computing nodes.

[0153] For in-network computing, the industry has proposed an implementation solution based on the Scalable Hierarchical Aggregation Protocol (SHARP). Specifically, an independently running management process is deployed on the management node, and this management process specifically includes a subnet manager (SM) and an aggregation manager (AM).

[0154] The SM obtains the topology information of the aggregation node (AN), then the SM informs the AM of this topology information. Next, the AM obtains the in-network computing capabilities of the AN, and the AM calculates the SHARP tree structure based on the AN topology information and the in-network computing capabilities of the AN. The AM allocates and configures the reliable connected queue pair between ANs, and then configures the SHARP tree information of all ANs according to the QP.

[0155] Next, the job scheduler starts the job and allocates computing resources. Each allocated host executes the job initialization script and starts the SHARP Deamon (SD). The SD numbered 0 (rank 0) (which can also be called SD-0) sends the job information to the AM, and the AM allocates SHARP resources for the job. The AM makes a quota for this job on the AN and sends the resource allocation description to SD-0, and SD-0 forwards the information to other SDs. It should be noted that other SDs can start the job in parallel, such as sending the job information to the AM, having the AM allocate resources for the job, and making a quota on the corresponding AN, etc.

[0156] In this way, the MPI process can access the SD to obtain SHARP resource information, establish a connection based on this SHARP resource information, then create a process group, and then the MPI process sends an aggregation request to the SHARP tree, thus realizing in-network computing.

[0157] The above method requires the deployment of management processes such as SM and AM on the management node, and obtains the network topology and INC resource information based on SM and AM, thereby realizing in-network computing. Among them, the deployment process of the management process is complex and difficult to maintain. When networking on a large scale, the difficulty of deploying and maintaining the management process is even more prominent.

[0158] In view of this, an embodiment of the present application provides a method for processing control messages in a collective communication system. The collective communication system includes a switch network and multiple computing nodes. Among them, the switch network includes at least one switch. The context of the collective communication system allows the communication space to be divided. Each context can provide a relatively independent communication space. Different messages can be transmitted in different contexts (specifically, different communication spaces), and the messages transmitted in one context will not be transmitted to another context. The computing nodes of the collective communication system can utilize the above characteristics of the context to initiate a control message process. Specifically, according to the existing context of the collective communication system, such as the context of the application or the context of the communication domain, a control message is generated and sent to other computing nodes of the collective communication system. Through this control message, the network topology and in-network computing capabilities of the switch network passed by the control message are queried for subsequent service messages to perform INC offloading operations.

[0159] This method reuses the context of the collective communication system to directly send and receive control messages such as query messages and notification messages, and queries the in-network computing capabilities based on the control messages, so that subsequent service messages can perform INC offloading based on the in-network computing capabilities, avoiding the repeated creation and acquisition of relevant resources, and decoupling the dependence on the control plane management process and the computing node daemon process. Based on this, the in-network computing solution provided by the embodiment of the present application is more maintainable, flexible, and general.

[0160] Furthermore, this method supports reusing existing network protocol channels, such as Ethernet channels, without relying on communication standards, without adopting remote direct memory access (RDMA) network communication standards (infiniband, IB), and without additional configuration of IB switches, thus greatly reducing the cost of the in-network computing solution.

[0161] Moreover, this method also does not require a daemon process to run on the computing node. Only an INC dynamic library (INC lib) needs to be provided, and the specified API in the INC lib is called within the collective operation communication domain to implement the control message service logic.

[0162] In order to make the technical solution of the present application clearer and easier to understand, the technical solution of the embodiment of the present application will be introduced below in combination with the system architecture diagram.

[0163] See Figure 1Architecture diagram of the collective communication system shown. The collective communication system 100 includes a switch network 102 and multiple computing nodes 104. The multiple computing nodes 104 include a master node and at least one child node. Among them, the master node is the computing node where the root process (the process with process number 0 in the ordered process series) in the collective communication system is located, and the child node is the computing node other than the master node in the collective communication system. A sub-root process may be included on the child node.

[0164] The switch network 102 includes at least one layer of switches 1020. Specifically, the switch network 102 can be a single-layer switch architecture. For example, it includes one layer of access switches, and the access switch can specifically be a top of rack (ToR) switch, thereby enabling the interconnection of computing nodes such as servers and the switch 1020 within the cabinet. It should be noted that the ToR switch can actually also be deployed at other positions within the cabinet, such as the middle of the cabinet, as long as it can achieve the interconnection of servers and switches within the cabinet.

[0165] The switch network 102 can also be a multi-layer switch architecture. For example, as Figure 1 shown, the switch network 102 can be a leaf-spine architecture. The switch network 102 with a leaf-spine architecture includes upper-layer switches, i.e., spine switches, located in the upper layer, and lower-layer switches, i.e., leaf switches, located in the lower layer. Among them, the spine switch is no longer the large chassis switch in the three-layer architecture but a switch with a high port density. The leaf switch can serve as the access layer. The leaf switch provides network connections to terminals and servers and is also connected to the spine switch upstream. It should be noted that the number of switches 1020 in each layer can be one or multiple, and the embodiments of the present application do not limit this.

[0166] The computing node 104 is a device with data processing capabilities, and can specifically be a server, or a terminal device such as a desktop computer, a laptop computer, or a smart phone. The multiple computing nodes 104 can be homogeneous devices. For example, the multiple computing nodes can all be servers with an Intel complex instruction set (X86) architecture, or all be servers with an advanced RISC machine (ARM) architecture. The multiple computing nodes 104 can also be heterogeneous devices. For example, some of the computing nodes 104 are servers with an X86 architecture, and some of the computing nodes are servers with an ARM architecture.

[0167] The switch network 102 and multiple computing nodes 104 form an HPC cluster. Any one or more of the computing nodes 104 in the cluster can serve as the storage nodes of the cluster. In some implementations, the cluster can also add independent nodes as the storage nodes of the cluster.

[0168] The switch network 102 is connected to each computing node 104 and acts as an in-network computing node. When the root process of the master node triggers a collective operation, such as a reduction summation operation, as Figure 2 shown, when the switch 1020 in the switch network 102, such as a leaf switch or a spine switch, receives a collective communication message, specifically a service message of the collective communication system, it can aggregate the service message and offload the calculation to the in-network computing engine (INC engine) of the switch. The INC engine performs in-network computing on the aggregated service message. Then, the switch 1020 forwards the calculation result to the computing node 104, which can reduce the load on the computing node 104. Among them, the calculation is jointly completed on the side of the switch network 102 and the computing node 104, reducing the number of transceiver times of the computing node 104, shortening the communication time, and improving the performance of collective communication.

[0169] It should be noted that for the switch network 102 to offload the calculation to the INC engine of the switch and for the INC engine to perform in-network computing on the aggregated message, the computing node 104 needs to know in advance the networking information (including the networking topology) and in-network computing capabilities of the in-network computing, and request corresponding resources according to the networking information and in-network computing capabilities. The networking information and in-network computing capabilities can be obtained through control messages.

[0170] Specifically, applications that support collective communication can be distributedly deployed on the computing nodes 104 of the collective communication system. When the application is initialized, a control message process can be initiated on the side of the computing node 104. Specifically, taking one of the computing nodes 104 in the collective communication system as the destination node and at least one of the remaining computing nodes 104 as the source node, the source node generates a query message according to the context of the collective communication system and transmits the above query message to the destination node.

[0171] Among them, when the query message passes through the switch 1020 in the switch network 102, the switch 1020 can add the in-network computing capabilities of the switch 1020 to the control message and then forward the above query message to the destination node. The destination node can generate a notification message according to the query message with the in-network computing capabilities of the switch 1020 added and return the above notification message to the source node, thereby notifying the source node of the in-network computing capabilities of the switch network 102. In this way, it is possible to obtain the in-network computing capabilities through control messages.

[0172] In some implementations, the network card in the computing node 104 also has certain computing capabilities. Based on this, the computing node 104 can also offload computing to the network card, thereby achieving in-node offloading. Specifically, when collective communication involves intra-node communication and inter-node communication, the intra-node computing can be offloaded to the network card, and the inter-node computing can be offloaded to the switch 1020. In this way, the performance of collective communication in a large-scale cluster can be further optimized.

[0173] Taking the collective communication among 32 processes located on 8 computing nodes 104 as an example, each computing node 104 includes 4 processes. The 4 processes can be aggregated on the network card in the computing node 104, and the computing of these 4 processes can be offloaded to the network card. The network card forwards the computing results of the 4 processes to the switch 1020, and the switch 1020 further aggregates the computing results of different computing nodes 104 and offloads the computing to the switch 1020. In this way, an in-network computing solution based on the network card and the switch 1020 can be implemented.

[0174] Multiple computing nodes in the collective communication system can be a primary node and at least one secondary node. In some possible implementations, the source node can be the above-mentioned secondary node, and the destination node can be the above-mentioned primary node, that is, the secondary node sends a query message to the primary node to query the in-network computing capabilities of the switch network 102. In some other possible implementations, the source node can also be the primary node, and the destination node can also be the secondary node, that is, the primary node sends a query message to the secondary node to query the in-network computing capabilities of the switch network 102.

[0175] The above introduced the architecture of the collective communication system. Next, the devices in the collective communication system, such as the switch 1020 and the computing node 104, will be introduced from the perspective of hardware implementation.

[0176] Figure 3 The structural schematic diagram of the switch 1020 is shown. It should be understood that Figure 3 Only some of the hardware structures and some software modules in the above-mentioned switch 1020 are shown. In specific implementation, the switch 1020 may further include more hardware structures, such as indicator lights, etc., and more software modules, such as various application programs, etc.

[0177] As Figure 3 shown, the switch 1020 includes a bus 1021, a processor 1022, a communication interface 1023, and a memory 1024. The processor 1022, the memory 1024, and the communication interface 1023 communicate with each other through the bus 1021.

[0178] The bus 1021 can be a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 only a thick line is used to represent it in Figure 3 , but it does not mean that there is only one bus or one type of bus.

[0179] The processor 1022 can be any one or more of processors such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP).

[0180] The communication interface 1023 is used for external communication, such as receiving query messages sent by sub-nodes and sending notification messages generated by the master node to sub-nodes, and so on.

[0181] The memory 1024 can include volatile memory, such as Random Access Memory (RAM). The memory 1024 can also include non-volatile memory, such as Read-Only Memory (ROM), flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD).

[0182] Programs or instructions are stored in the memory 1024, such as the programs or instructions required to implement the control message processing method in the set communication system provided in the embodiments of the present application. The processor 1022 executes the program or instruction to execute the foregoing control message processing method in the set communication system.

[0183] It should be noted that Figure 3Only one switch 1020 in the switch network 102 is shown. In some implementations, the switch network 102 may include multiple switches 1020. Considering the transmission performance between the switches 1020, these multiple switches 1020 may also be integrated on a backplane or placed on the same rack.

[0184] Figure 4 A schematic structural diagram of the computing node 104 is shown. It should be understood that Figure 4 Only some of the hardware structures and some of the software modules in the above computing node 104 are shown. In specific implementations, the computing node 104 may further include more hardware structures, such as microphones, speakers, etc., and more software modules, such as various application programs, etc.

[0185] Such as Figure 4 As shown, the computing node 104 includes a bus 1041, a processor 1042, a communication interface 1043, and a memory 1044. The processor 1042, the memory 1044, and the communication interface 1043 communicate with each other through the bus 1041.

[0186] The bus 1041 may be a PCI bus, PCIe, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of easy representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The processor 1042 may be any one or more of processors such as a CPU, a GPU, an MP, or a DSP. The communication interface 1043 is used for external communication, for example, transmitting query messages through the switch network 102, transmitting notification messages through the switch network 102, and so on.

[0187] The memory 1044 may include volatile memory, such as random access memory. The memory 1044 may also include non-volatile memory, such as read-only memory, flash memory, a hard disk drive, or a solid-state drive. Programs or instructions are stored in the memory 1044, such as the programs or instructions required to implement the control message processing method in the set communication system provided by the embodiments of the present application. The processor 1042 executes the program or instruction to execute the foregoing control message processing method in the set communication system.

[0188] In order to make the technical solutions of the present application clearer and easier to understand, the control message processing method in the set communication system provided by the embodiments of the present application will be introduced in detail below with reference to the accompanying drawings.

[0189] See Figure 5 The control message processing method in the set communication system shown, the method includes:

[0190] S502: The switch 1020 forwards the query message transmitted from the source node to the destination node.

[0191] To simplify the communication process of service messages in the collective communication system and optimize the performance of collective communication, the computing node 104 may first transmit a control message, obtain information such as the in-network computing ability based on the control message interaction on the control plane, so as to lay a foundation for the transmission of service messages.

[0192] In some possible implementation manners, a child node in the computing node 104 may initiate a control message process as a source node. Specifically, the child node (specifically, the sub-root process of the child node) may generate a query message according to the context of the collective communication system, such as the context of the application or the context of the communication domain. The query message is a type of control message and is used to query the in-network computing ability of the switch network 102. The in-network computing ability may also be referred to as the computing offloading ability and is used to characterize the ability of the switch network 102 to undertake computing tasks. Among them, the in-network computing ability may be characterized by at least one of the following indicators: the type of collective operation supported, the data type.

[0193] Among them, the type of collective operation supported by the switch network 102 may include any one or more of the following: broadcast from one member to all members in the group, collect data from all members by one member, scatter data from one member to all members in the group, scatter / collect data operation between all members in the group, global reduction operation, combined reduction and scatter operation, search operation on all members in the group. Among them, a member refers to a process in the process group.

[0194] The data types supported by the switch network 102 may include any one or more of byte, 16-bit integer (short), 32-bit integer (int), 64-bit integer (long), floating point type (float), double precision floating point type (double), boolean type (boolean), and character type (char), etc.

[0195] The query message is transmitted from the child node to the master node through the switch network 102. Specifically, the query message is generated according to the context of the collective communication system, such as the context of the communication domain. The query message may carry a communication domain identifier. For example, the communication domain identifier is carried in the INC message header. The switch 1020 forwards the above query message to the destination node within the communication domain based on the communication domain identifier. The child node can obtain the in-network computing power of the switch network 102 according to the notification message corresponding to the above query message, without obtaining the networking information (specifically, the topology information of the switch network 102) through the SM process on the management node, and informing the AM of the topology information, and the AM obtains the in-network computing power of the switch network 102.

[0196] In some implementation manners, the switch 1020 (which can also be referred to as the first switch) that receives the query message in the switch network 102 may add the in-network computing power of the switch 1020 to the query message, and then forward the query message with the in-network computing power of the above switch 1020 added to the master node (specifically, the root process of the master node).

[0197] Specifically, the query message includes an INC message header. The INC message header is used to identify that the message is an INC message, including an INC control message or an INC data message (also referred to as an INC service message). When the switch 1020 receives a message, it can identify whether the message is an INC message according to whether it includes the INC message header, and then decide whether to perform the operation of querying the in-network computing power.

[0198] To query the in-network computing power, a query field may be reserved in the query message. The switch 1020 adds the in-network computing power of the switch 1020 to the query field. The in-network computing power of the switch 1020 includes in-network computing characteristics such as the operation types and data types supported by the switch 1020. These in-network computing characteristics can form an in-network computing characteristic list (INC feature list). The switch 1020 can add the INC feature list to the query field.

[0199] For ease of understanding, an example of the query message is also provided in an embodiment of the present application. The query message includes an INC message header and the payload of the INC message. As Figure 6As shown, the INC message header is the MPI+ field, and the payload of the INC message is the query field. In some implementations, the header of the query message further includes an MPI message header, an IB message header, a user datagram protocol (UDP) message header, an Internet protocol (IP) message header, an Ether message header, etc., for transmission in the transport layer and Ethernet. Of course, the tail of the query message may further include a check field, such as a cyclic redundancy check (CRC).

[0200] Among them, the INC message header may specifically include an in-network computing tag (1NC Tag) and a communicator ID (commID). Among them, the in-network computing tag includes INC tag low. In some examples, when the value of INC tag low is 0x44332211, it indicates that the message is an INC message. The in-network computing tag may further include INC tag high, and INC tag high and INC tag low can be jointly used to identify the message as an INC message, thereby improving the accuracy of identifying INC messages. The communicator ID specifically identifies the communication domain multiplexed by this collective communication. In some possible implementations, the INC message header further includes the operation type (operation code, opt code) and data type of this collective communication.

[0201] The INC message header may further include a source rank (src rank), that is, the rank number of the process that generates this message. The INC message header may further include one or more of a request ID (req ID), a request packet number (ReqPkt Num), and a PktPara Num of the message. In some embodiments, the INC message header may further include a reserved field, such as a reserve ID (rsvd), and different reserved values of this reserve ID can be used to distinguish control messages sent by computing nodes and control messages sent by switches.

[0202] The query field includes the supported data operations, supported data types that the switch supports. In some implementations, the query field may further include the supported MPI type, supported coll type, maximum data size, global group size, local group size, communication domain ID, available group number. Among them, the query field may further include querynotify hop. Query notify hop may occupy one byte. When the first 4 bits of this byte take the value of 0x0, it indicates that this message is a query message.

[0203] The query message passes through a switch 1020, and the switch 1020 fills in its in-network computing capability in the query message (specifically, the query field of the query message). When the query message passes through multiple switches 1020, each switch 1020 adds its in-network computing capability to the query message, specifically adding the supported operation type and / or data type of the switch 1020.

[0204] Furthermore, the in-network computing capability of the switch 1020 may also include the size of the remaining available resources for in-network computing of the switch 1020. Among them, the size of the remaining available resources for in-network computing of the switch 1020 can be characterized by the maximum number of concurrent hosts of the switch 1020, also known as the local group size. Based on this, the switch 1020 can also add the local group size to the query message. In some embodiments, the switch 1020 can also add any one or more of the supported MPI type, supported coll type, maximum data size, global group size, communication domain ID, etc.

[0205] The master node (specifically, the root process on the master node) can summarize the in-network computing capabilities of these switches 1020 to obtain the in-network computing capability of the switch network 102, and generate a notification message based on the in-network computing capability of the switch network 102 to notify the sub-root process of the child node of the in-network computing capability of the switch network 102.

[0206] In some implementations, the switch 1020 can also fill in the message hop count (hop) in the query field. See Figure 6, switch 1020 can add the message hop count in the last 4 bits of the query notify hop. Correspondingly, switch 1020 can also create an entry based on the message hop count.

[0207] Specifically, when the message hop count is 0, it indicates that the source node is directly connected to this switch 1020, and this switch 1020 creates an entry within the switch. In this way, in the service message process, when switch 1020 receives a service message, it can perform computational offloading on the service message according to the above-mentioned entry.

[0208] In some embodiments, the entry at least includes the identifier of the source node (specifically, the process on the source node), such as source rank. Switch 1020 first identifies the service message as an INC message according to the message header, and then compares the srcrank in the message with the src rank in the entry. When the src rank is the same, switch 1020 allocates INC resources for computational offloading. Further, when the collective communication is completed, switch 1020 can also delete the above-mentioned entry and release the INC resources. Thus, real-time allocation and release of INC resources are achieved, optimizing resource utilization.

[0209] S504: The switch 1020 forwards the notification message transmitted from the destination node to the source node.

[0210] The notification message is used to notify the source node of the in-network computing ability of the switch network 102. The notification message is generated by the destination node according to the query message. Similar to the query message, the context of the collective communication system, such as the communication domain identifier, is carried in the notification message. In addition, the in-network computing ability of the switch network 102 is also carried in the notification message. Among them, the context enables the notification message to be accurately transmitted to the corresponding node (specifically, the process of the node, such as the sub root process of the sub node). In this way, source nodes such as sub nodes can obtain the in-network computing ability of the switch network 102 without starting the SM process and AM process on the management node.

[0211] For ease of understanding, Figure 7 a specific example of the notification message is also provided. As Figure 7 shown, the format of the notification message is the same as that of the query message. The in-network computing ability of the switch network 102, including supported data operation, supported data type, etc., is filled in the query field of the notification message. Among them, the query field can also include query notify hop. Query notify hop can occupy one byte. When the first 4 bits of this byte take the value of 0x1, it indicates that this message is a notification message.

[0212] When the switch network 102 includes a single-layer switch 1020, specifically an access-layer switch, the single-layer switch can forward the notification message to the child node when receiving the notification message, so as to realize the transmission of the notification message generated by the master node to the child node.

[0213] When the switch network 102 includes a multi-layer switch 1020, specifically an upper-layer switch and a lower-layer switch, the lower-layer switch can determine the target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing. The target switch is specifically used to aggregate service messages in the subsequent service message process, thereby realizing in-network computing. Thus, the lower-layer switch can add the size of the remaining available resources for in-network computing of the target switch to the notification message. The size of the remaining available resources for in-network computing of the target switch can be characterized by the maximum number of concurrent hosts of the target switch, that is, the available group size. Based on this, the lower-layer switch can add the available group size to the query field of the notification message.

[0214] The in-network computing ability of the switch network 102 can also include the available group size, that is, the size of the remaining available resources for in-network computing of the target switch. The lower-layer switch can forward the notification message with the size of the remaining available resources for in-network computing of the target switch added to the child node. Correspondingly, the child node can initiate a service message process according to the in-network computing ability of the switch network 102 in the notification message to realize in-network computing.

[0215] Among them, when the lower-layer switch determines the target switch, it specifically determines the target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing by using a load balancing strategy. For example, the switch network 102 includes n switches, where m switches are upper-layer switches, and n is greater than m. When the master node returns a notification message to the child node, the notification message is forwarded through the switch network 102. When the notification message reaches the lower-layer switch close to the child node, the lower-layer switch selects the switch with larger available resources for in-network computing and smaller load among the m upper-layer switches as the target switch according to the size of the available resources for in-network computing of the upper-layer switches through a load balancing strategy.

[0216] In some implementation manners, the lower-layer switch can send a switch query message to the upper-layer switch, and the switch query message is used to query the size of the remaining available resources for in-network computing. Then, the upper-layer switch can return a switch notification message to the lower-layer switch, and the switch notification message carries the size of the remaining available resources for in-network computing.

[0217] After determining the target switch, the lower-layer switch may also send a resource request message to the target switch. The resource request message is used to request resource allocation. The size of the requested allocated resource does not exceed the size of the remaining available resources of the target switch. The target switch allocates resources according to the resource request message. After successful allocation, the target switch may establish an entry so that in subsequent service message processes, the subsequent service messages can be aggregated through the allocated resources and the corresponding entries, thereby realizing computing offloading. The target switch may also generate a resource response message and send the resource response message to the above-mentioned lower-layer switch. The resource response message is used to notify the above-mentioned lower-layer switch that the resource allocation is successful.

[0218] It should be noted that Figure 5 In the illustrated embodiment, the sub-node generates a query message and sends the query message to the master node. The master node generates a notification message according to the query message with the online computing power of switch 1020 added, and returns the notification message to the sub-node for illustrative purposes. In other possible implementation manners of the embodiments of the present application, it may also be that the master node generates a query message and the sub-node generates a notification message, or one sub-node generates a query message and another sub-node generates a notification message. The embodiments of the present application do not limit this.

[0219] In Figure 5 In the illustrated embodiment, when switch 1020 is directly connected to the source node and the destination node, switch 1020 may directly receive the query message sent by the source node, forward the query message to the destination node, then receive the notification message sent by the destination node, and forward the notification message to the source node. In this way, the sending and receiving of the query message and the notification message can be realized through one forwarding, improving the efficiency of obtaining the online computing power.

[0220] When switch 1020 (also referred to as the first switch) is connected to the source node through other switches (for the sake of description, referred to as the second switch in the present application), or is connected to the destination node through other switches (for the sake of description, referred to as the third switch in the present application), switch 1020 may also forward the message to other switches, and forward the message to the source node or the destination node through other switches.

[0221] Specifically, when the switch network 102 includes a second switch and does not include a third switch, switch 1020 receives the query message forwarded by the second switch, forwards the query message to the destination node, and then switch 1020 receives the notification message sent by the destination node and forwards the notification message to the second switch.

[0222] When the switch network 102 includes a third switch and does not include a second switch, the switch 1020 receives the query message sent by the source node, forwards the query message to the third switch, and then the switch 1020 receives the notification message forwarded by the third switch and forwards the notification message to the source node.

[0223] When the switch network 102 includes both a second switch and a third switch, the switch 1020 receives the query message forwarded by the second switch, forwards the query message to the third switch, and then the switch 1020 receives the notification message forwarded by the third switch and forwards the notification message to the second switch.

[0224] The upper-layer switch and the lower-layer switch may include switches of more than one level, that is, the multi-layer switch may be a switch of 2 layers or more than 2 layers. For the sake of description, in the embodiments of the present application, the multi-layer switch included in the switch network 102 is also taken as a leaf-spine architecture switch to exemplarily illustrate the method for processing control messages in the collective communication system.

[0225] See Figure 8 The flowchart of the method for processing control messages in the collective communication system shown. In this example, the collective communication system includes a main node, at least one sub-node (such as sub-node 1 and sub-node 2) and a switch network 102, where the switch network 102 includes leaf switches and spine switches. This example is described with 2 leaf switches (specifically leaf1 and leaf2) and 2 spine switches (specifically spine1 and spine2). The method includes:

[0226] S802: The sub-node generates a query message according to the context of the communication domain in the collective communication system and sends the query message to the main node.

[0227] Specifically, the sub-root process on the sub-node sends a query message to the root process on the main node to query the in-network computing ability of the switch network 102. The query message is generated by the sub-node according to the context of the communication domain, so as to ensure that the query message can be correctly transmitted to the main node. Among them, the sub-node may be a server. For the sake of distinguishing it from other query messages, the query message sent by the sub-node may be called a server query.

[0228] Among them, the collective communication system includes multiple sub-nodes. For example, when the collective communication system includes sub-node 1 and sub-node 2, the query message sent by sub-node 1 is specifically server1 query, and the query message sent by sub-node 2 is specifically server2 query.

[0229] In a collective communication system, there is a communication domain. For example, the communication domain corresponding to the process group formed by the root process, sub-root process 1, and sub-root process 2. The child nodes (specifically, the sub-root processes on the child nodes, such as sub-root process 1 and sub-root process 2) reuse the context of the communication domain and send a serverquery to the master node (specifically, the root process on the master node), without having to deploy a management node and start management processes such as SM and AM on the management node.

[0230] In Figure 8 the example of, since server1 and server2 are directly connected to different switches respectively. For example, server1 is directly connected to leaf1, server2 is directly connected to leaf2, and leaf2 is also directly connected to the master node while leaf1 is not directly connected to the master node. Therefore, the server1query needs to pass through leaf1, spine1, and leaf2 to reach the master node, and the server2 query passes through leaf2 to reach the master node, and their paths are different.

[0231] S804: When the switch 1020 in the switch network 102 receives a query message, the in-network computing ability of this switch 1020 is added to the query message.

[0232] The query field is reserved in the server query (such as server query1 and server query2). When the serverquery is transmitted to the master node through the switch, the switch 1020 adds the in-network computing ability of this switch 1020 to the query field of the server query. Among them, when the server query passes through multiple switches 1020, the in-network computing abilities of these multiple switches 1020 are added to the query field.

[0233] Among them, the in-network computing ability of the switch 1020 specifically includes any one or more of the supported collective operation types, data types, etc. Further, the in-network computing ability of the switch 1020 also includes the size of the remaining available resources of the in-network computing of this switch 1020.

[0234] In some implementation manners, the query field is also used to add the message hop count (hop). The switch 1020 can also create an entry according to the communication domain identifier and the message hop count. Specifically, when the message hop count is 0, it indicates that this switch 1020 is directly connected to the child node, and the switch 1020 can create an entry. In the subsequent service message process, the switch 1020 aggregates the service messages according to this entry, thereby realizing computing offloading.

[0235] S806: The master node generates a notification message and sends the notification message to the slave nodes.

[0236] Specifically, the master node can summarize the field values of the query field in the received server query (specifically, the server query with the in-network computing power of switch 1020 added) to obtain the in-network computing power of switch network 102. In addition, the master node can also obtain the networking information of switch network 102, such as topology information, based on the switch forwarding path.

[0237] The master node can generate a notification message based on information such as the in-network computing power and return the notification message to the corresponding slave nodes. For the sake of easy distinction, this notification message can be called server notify. When multiple slave nodes send server queries, the master node can return the corresponding notification messages, such as server1 notify and server2 notify. In Figure 8 the example, server1 and serve2 are respectively connected to different switches. Therefore, the paths of server1 notify and server2 notify are different.

[0238] S808: When the lower-layer switch (access layer switch) close to the slave node receives the notification message, it sends a switch query message to the upper-layer switch.

[0239] The lower-layer switch close to slave node 1 is leaf 1, and the lower-layer switch of slave node 2 is leaf2. When leaf 1 receives server1 notify, leaf 1 sends a switch query message switch query to the upper-layer switches (specifically, spine1 and spine2). When leaf 2 receives server2 notify, leaf 2 sends switch query to spine1 and spine2. This switch query is used to query the size of the remaining available resources of the in-network computing.

[0240] S810: The upper-layer switch returns a switch notification message to the lower-layer switch.

[0241] For the sake of easy description, the switch notification message can be called switch notify. Switch notify is used to notify the lower-layer switch of the size of the available resources of the in-network computing of the upper-layer switch.

[0242] S812: The lower-layer switch determines the target switch according to the size of the available resources of the in-network computing and sends a resource request message to the target switch.

[0243] Specifically, lower-layer switches such as leaf 1 and leaf 2 aggregate the resource information feedback by each upper-layer switch through switch notify, and determine the target switch according to the remaining available resources calculated by each upper-layer switch in the network. When determining the target switch, the lower-layer switch can determine the target switch through a load balancing strategy. This can balance the load of each switch and increase the number of concurrent collective communications.

[0244] In some implementation manners, the lower-layer switch can also randomly select a switch from the switches whose remaining available resources are greater than or equal to the requested resources as the target switch. The embodiments of the present application do not limit the implementation manner of determining the target switch.

[0245] After the lower-layer switch determines the target switch, it can send a resource request message to the target switch to request resource allocation. In Figure 8 In the illustrated embodiment, the target switch is spine1, and leaf 1 and leaf 2 respectively send resource request messages to spine1. The resource request message is a request message sent by the lower-layer switch to the target switch, and thus can be denoted as switch request for distinction.

[0246] S814: The target switch sends a resource response message to the lower-layer switch.

[0247] Specifically, the target switch such as spine1 can count whether the switch requests in the communication domain of the current collective communication are received completely. After the switch requests are received completely, it can allocate resources and then return a resource response message to the lower-layer switch. Similar to the switch request, the resource response message can be denoted as switch response.

[0248] Among them, the switch request includes a global group size field and a local group size field. Among them, the global group size is also called the global host number, and the local group size is also called the local host number. The target switch such as spine1 can determine whether the switch requests are received completely according to the field values of the local host number field and the global host number field. Specifically, the target switch can aggregate the switch requests, sum up the local host numbers, and then compare the sum of the local host numbers with the global host number. If they are equal, it indicates that the switch requests are received completely. If not, it will judge whether the request messages are received completely according to the local host number and the global host number fields in the INC message.

[0249] Further, when the target switch receives a switch request from the lower-layer switch, it can also create an entry, which is used to allocate resources for subsequent service packets, aggregate the service packets, and thus achieve computing offloading.

[0250] In some embodiments, when executing the control packet processing method in the collective communication system according to the embodiments of the present application, the steps of the lower-layer switch sending a resource request packet to the target switch in S812 and the target switch sending a resource response packet to the lower-layer switch in S814 may not be executed. After determining the target switch, the lower-layer switch can directly execute S816, and the process of requesting resource allocation can be executed in the subsequent service packet process.

[0251] S816: The lower-layer switch adds the size of the remaining available in-network computing resources of the target switch to the notification packet, and forwards the notification packet with the size of the remaining available in-network computing resources of the target switch added to the child node.

[0252] This notification packet is a server notify. The switch network 102 includes upper-layer switches, such as spine1 and spine 2. The service packets of leaf 1 and leaf2 can be aggregated at spine 1 or spine2. When the lower-layer switch determines that spine 1 is the target switch and successfully requests the target switch to allocate resources, it can also add the size of the remaining available in-network computing resources of the target switch spine1 to the server notify, that is, add the available group size to the query field. Then the lower-layer switch sends the server notify with the available group size added to the child node. This server notify is used to notify the child node of the in-network computing capabilities of the switch network 102, such as the collective operation types, data types supported by the switch 1020, and the size of the remaining available in-network computing resources of the target switch, so as to lay a foundation for the in-network computing of subsequent service packets. Among them, leaf 1 sends server1 notify to child node 1, and leaf 2 sends server2 notify to child node 2.

[0253] When the switch network includes only one spine, since the service packets reported by the leaf can be directly aggregated based on this spine without performing a selection operation. Therefore, in the control packet process, when the lower-layer switch directly connected to the source node receives the notification packet, it may not execute S808 to S816 above, but directly create an entry on this spine and forward the notification packet to the source node.

[0254] Figure 8The illustrated embodiments are mainly exemplified by a switch network 102 including leaf and spine switches. In some implementations, the switch network 102 includes a single-layer switch, which is an access-layer switch. The single-layer switch may include one or more switches, such as one or more Tors. Hereinafter, the switch network 102 including Tor1 and Tor2 is used for exemplary illustration.

[0255] See Figure 9 The flowchart of the control message processing method in the illustrated collective communication system. The method includes:

[0256] S902: The child node generates a query message according to the context of the communication domain in the collective communication system and sends the query message to the master node.

[0257] S904: When the switch 1020 of the switch network 102 receives the query message, the in-network computing power of the switch 1020 is added to the query message.

[0258] For the specific implementation of S902 to S904, reference can be made to the relevant content description of S602 to S604, which will not be elaborated herein in the embodiments of the present application.

[0259] S906: The master node generates a notification message according to the query message added with the in-network computing power of the switch 1020 and sends the notification message to the child node.

[0260] Among them, the master node sends the notification message to the child node through the switch network 102. The master node does not need to send a switch query message to the upper-layer switch through the switch in the switch network 102 to determine the size of the available remaining in-network computing resources of the upper-layer switch, nor send a resource request message to the upper-layer switch so that the upper-layer switch converges the service messages. Instead, the master node directly aggregates the service messages on Tor1 and Tor2 to achieve in-network computing.

[0261] In Figure 6 、 Figure 7 In the illustrated embodiments, the computing node 104 may initiate the control message process in a polling manner. In this way, when the topology of the switch network 102 changes, the computing node 104 can obtain the in-network computing power of the switch network 102 in real time. In some implementations, the computing node 104 may also initiate the control message process periodically to update the in-network computing power of the switch network 102 in a timely manner when the topology of the switch network 102 changes.

[0262] Figure 8 、 Figure 9The method for processing control messages in the collective communication system provided by the embodiments of the present application will be exemplarily described from the perspectives of the switch network 102 including a multi-layer switch and the switch network 102 including a single-layer switch respectively.

[0263] Next, the method for processing control messages in the collective communication system provided by the embodiments of the present application will be described in conjunction with a specific scenario, such as a weather prediction scenario.

[0264] In the weather prediction scenario, by building an HPC cluster and deploying a weather research and forecasting model (WRF) on the HPC cluster, fine-scale weather simulation and forecasting can be achieved.

[0265] Specifically, the HPC cluster may include a switch 1020 and eight computing nodes 104. Among them, the switch 1020 may be a 10-gigabit (G) Ethernet switch. The switch 1020 includes processors, such as one or more of a central processing unit (CPU) and a neural-network processing unit (NPU). The switch 1020 realizes in-network computing through the above CPU and NPU, reducing the computing pressure on the computing nodes 104. The computing nodes 104 may be servers configured with 10G Ethernet network cards.

[0266] A user (such as an operation and maintenance personnel) deploys a community enterprise operating system (cent OS) on the server. Then, WRF is deployed on the operating system. When deploying WRF, it is necessary to first create an environment variable configuration file and install the dependency packages of WRF, such as hierarchical data format version 5 (HDF5), parallel network common data form (PnetCDF), and the dependency packages corresponding to different languages in netCDF, such as netCDF-C and netCDF-fortran. Then, the operation and maintenance personnel install the main program, that is, install the source code package of WRF. Among them, before installing the source code package, the validity of the environment variables can be determined first to ensure the normal operation of WRF.

[0267] One of the eight servers serves as the master node, and the remaining servers serve as slave nodes. Processes on the slave nodes (specifically, the processes of WRF) generate query messages according to the context of the communication domain in the collective communication system, and then send the query messages to the processes on the master node. When the query message passes through switch 1020, switch 1020 adds its in-network computing power to the query field of the query message, and then forwards the query message with the added in-network computing power to the master node. In this way, the process on the master node receives the query message with the in-network computing power of switch 1020 added, and obtains the in-network computing power of switch network 102 according to the query field of the query message, specifically, the in-network computing power of switch network 102 involved in this collective communication. Among them, the switch network 102 in this embodiment includes one switch 1020. Therefore, the in-network computing power of switch network 102 is the in-network computing power of switch 1020.

[0268] The process on the master node generates a notification message according to the above in-network computing power. The notification message is used to notify the in-network computing power of switch network 102. Then, when switch 1020 receives the notification message, it forwards the above notification message to the process on the slave node.

[0269] In this way, the process on the slave node can know the in-network computing power of switch network 102. When the process on the slave node performs collective communication, such as executing a broadcast operation from one member to all members in the group, the slave node can also offload the calculation to switch 1020 according to the in-network computing power. Thus, the in-network computing scheme of WRC is realized, and the efficiency of weather forecasting is improved.

[0270] It should be noted that the above switch 1020 can support INC in hardware and support processing query messages in software, specifically, adding the in-network computing power of the switch 1020 to the query message. The switch 1020 can be a self-developed switch with the above functions, or a switch obtained by transforming an existing switch based on the above method provided in this application embodiment. The computing node 104 can be a self-developed server or a general-purpose server, and the corresponding MPI is deployed on the server.

[0271] The above method provided by the embodiments of the present application can also be applied to a cloud environment. Specifically, the computing node 104 can be a cloud computing device in a cloud platform. For example, the computing node 104 can be a cloud server in an Infrastructure as a Service (IaaS) platform. The switch 1020 in the switch network 102 can be a switch in the cloud platform, that is, a cloud switch. The cloud computing device multiplexes the context of the communication domain to transmit control messages, thereby obtaining the in-network computing ability of the cloud switch. According to this in-network computing ability, the load is offloaded to the cloud switch, which can optimize the collective communication performance and provide a more flexible and efficient on-demand allocation service.

[0272] As described above in conjunction with Figures 1 to 9 The method for processing control messages in the collective communication system provided by the embodiments of the present application has been introduced in detail. Next, the control message device, as well as devices such as switches, first computing nodes, and second computing nodes, provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.

[0273] See Figure 10 The structural schematic diagram of the control message processing device in the collective communication system shown. The collective communication system includes a switch network and multiple computing nodes. The switch network includes a first switch. The device 1000 includes:

[0274] A communication module 1002, configured to forward a query message transmitted from a source node to a destination node. The query message is used to request to query the in-network computing ability of the switch network. The query message is generated by the source node according to the context of the collective communication system. The source node and the destination node are different nodes among the multiple computing nodes;

[0275] The communication module 1002 is further configured to forward a notification message transmitted from the destination node to the source node. The notification message carries the in-network computing ability of the switch network.

[0276] In some possible implementation manners, the device 1000 further includes:

[0277] A processing module 1004, configured to add the in-network computing ability of the first switch to the query message when receiving the query message;

[0278] The communication module 1002 is specifically configured to:

[0279] Forward the query message with the in-network computing ability of the first switch added to the destination node.

[0280] In some possible implementation manners, the in-network computing ability of the first switch includes the type of collective operation and / or data type supported by the first switch.

[0281] In some possible implementations, the in-network computing capability of the first switch includes the size of the remaining available resources for in-network computing of the first switch.

[0282] In some possible implementations, the apparatus further includes:

[0283] A processing module 1004, configured to establish an entry according to the hop count of the query message, where the entry is used for the first switch to perform computing offloading on service messages.

[0284] In some possible implementations, the first switch is directly connected to the source node and the destination node;

[0285] The communication module 1002 is specifically configured to:

[0286] Receive a query message sent by the source node, and forward the query message to the destination node;

[0287] Receive a notification message sent by the destination node, and forward the notification message to the source node.

[0288] In some possible implementations, the switch network further includes a second switch and / or a third switch, where the second switch is used to connect the first switch and the source node, and the third switch is used to connect the first switch and the destination node;

[0289] The communication module 1002 is specifically configured to:

[0290] Receive a query message sent by the source node, and forward the query message to the third switch;

[0291] Receive a notification message forwarded by the third switch, and forward the notification message to the source node; or,

[0292] Receive a query message forwarded by the second switch, and forward the query message to the destination node;

[0293] Receive a notification message sent by the destination node, and forward the notification message to the second switch; or,

[0294] Receive a query message forwarded by the second switch, and forward the query message to the third switch;

[0295] Receive a notification message forwarded by the third switch, and forward the notification message to the second switch.

[0296] In some possible implementations, the switch network includes a single-layer switch, and the first switch is the single-layer switch;

[0297] The communication module 1002 is specifically configured to:

[0298] Forward the notification message to the source node.

[0299] In some possible implementation manners, the switch network includes an upper-layer switch and a lower-layer switch, and the first switch is the lower-layer switch;

[0300] The apparatus 1000 further includes:

[0301] A processing module 1004, configured to determine a target switch from the upper-layer switches according to the size of the remaining available resources calculated in the network of the upper-layer switch, and add the size of the remaining available resources calculated in the network of the target switch to the notification message;

[0302] The communication module 1002 is specifically configured to:

[0303] Forward the notification message added with the size of the remaining available resources calculated in the network of the target switch to the source node.

[0304] In some possible implementation manners, the processing module 1004 is specifically configured to:

[0305] Determine a target switch from the upper-layer switches according to the size of the remaining available resources calculated in the network, by using a load balancing strategy.

[0306] In some possible implementation manners, the communication module 1002 is further configured to:

[0307] Send a switch query message to the upper-layer switch, where the switch query message is used to query the size of the remaining available resources calculated in the network of the upper-layer switch;

[0308] Receive a switch notification message sent by the upper-layer switch, where the switch notification message is used to notify the size of the remaining available resources calculated in the network of the upper-layer switch.

[0309] In some possible implementation manners, the switch network includes an upper-layer switch and a lower-layer switch, and the first switch is the upper-layer switch;

[0310] The communication module 1002 is further configured to:

[0311] Receive a switch query message sent by the lower-layer switch, where the switch query message is used to query the size of the remaining available resources calculated in the network of the first switch;

[0312] Send a switch notification message to the lower-layer switch, where the switch notification message is used to notify the lower-layer switch of the size of the remaining available resources for in-network computing of the first switch.

[0313] In some possible implementation manners, the context of the collective communication system includes the context of an application or the context of a communication domain.

[0314] In some possible implementation manners, the multiple computing nodes include a master node and at least one slave node;

[0315] The source node is the slave node, and the destination node is the master node; or,

[0316] The source node is the master node, and the destination node is the slave node.

[0317] In the collective communication system according to the embodiment of the present application, the control message processing device 1000 may correspond to executing the method described in the embodiment of the present application, and the above and other operations and / or functions of each module / unit of the control message processing device 1000 in the collective communication system are respectively for implementing Figures 5 to 9 the corresponding processes of the respective methods in the illustrated embodiments. For the sake of brevity, they will not be elaborated here.

[0318] Figure 10 The control message processing device 1000 in the collective communication system provided by the illustrated embodiment is specifically a device corresponding to the switch 1020. The embodiment of the present application also provides devices corresponding to the first computing node and the second computing node respectively.

[0319] See Figure 11 the structural schematic diagram of the control message processing device 1100 in the illustrated collective communication system. The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The device 1100 includes:

[0320] A communication module 1102, configured to receive query messages forwarded by one or more switches in the switch network. The query messages are used to request to query the in-network computing capabilities of the switch network. The query messages are generated by the second computing node according to the context of the collective communication system;

[0321] A generation module 1104, configured to generate a notification message according to the query messages, where the notification message carries the in-network computing capabilities of the switch network;

[0322] The communication module 1102 is further configured to send the notification message to the second computing node.

[0323] In some possible implementation manners, the query message forwarded by the switch includes the in-network computing capabilities of the switch added by the switch.

[0324] The generating module 1104 is specifically configured to:

[0325] Obtain the in-network computing capabilities of the switch network according to the in-network computing capabilities of the one or more switches in the query messages forwarded by the one or more switches.

[0326] Generate a notification message according to the in-network computing capabilities of the switch network.

[0327] In some possible implementation manners, the device 1100 is deployed on the first computing node, and the first computing node is a master node or a slave node.

[0328] In the collective communication system according to the embodiments of the present application, the control message processing device 1100 may correspondingly execute the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module / unit of the control message processing device 1100 in the collective communication system are respectively for implementing Figure 8 or Figure 9 the corresponding processes of the respective methods in the illustrated embodiments. For the sake of brevity, they are not described herein again.

[0329] Next, refer to Figure 12 the structural schematic diagram of the control message processing device 1200 in the collective communication system shown. The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The device 1200 includes:

[0330] A generating module 1202, configured to generate a query message according to the context of the collective communication system, where the query message is used to request to query the in-network computing capabilities of the switch network.

[0331] A communication module 1204, configured to send the query message to the first computing node through one or more switches in the switch network.

[0332] The communication module is further configured to receive a notification message forwarded by the first computing node through the one or more switches. The notification message carries the in-network computing capabilities of the switch network, and the notification message is generated by the first computing node according to the query message.

[0333] In some possible implementation manners, the device is deployed on the first computing node, and the second computing node is a master node or a slave node.

[0334] In the control message processing apparatus 1200 in the collective communication system according to an embodiment of the present application, it may correspond to executing the method described in the embodiment of the present application, and the above and other operations and / or functions of each module / unit of the control message processing apparatus 1200 in the collective communication system are respectively for realizing Figure 8 or Figure 9 the corresponding processes of the respective methods in the embodiments shown. For the sake of brevity, they will not be elaborated here.

[0335] Based on Figure 10 、 Figure 11 、 Figure 12 the control message processing apparatus 1000 in the collective communication system, the control message processing apparatus 1100 in the collective communication system, and the control message processing apparatus 1200 in the collective communication system provided by the embodiments shown, an embodiment of the present application further provides a collective communication system 100.

[0336] For the sake of convenience of description, in the embodiment of the present application, the control message processing apparatus 1000 in the collective communication system, the control message processing apparatus 1100 in the collective communication system, and the control message processing apparatus 1200 in the collective communication system are respectively abbreviated as the control message processing apparatus 1000, the control message processing apparatus 1100, and the control message processing apparatus 1200.

[0337] Referring to Figure 13 the structural schematic diagram of the collective communication system 100 shown, the collective communication system 100 includes a switch network 102 and a plurality of computing nodes 104. Among them, the switch network 102 includes at least one switch 1020, and the switch 1020 is specifically used to implement the functions of the control message processing apparatus 1000 as shown in Figure 10 . A plurality of computing nodes 104 include a destination node and at least one source node. The destination node is specifically used to implement the functions of the control message processing apparatus 1100 as shown in Figure 11 . The source node is specifically used to implement the functions of the control message processing apparatus 1200 as shown in Figure 12 .

[0338] Specifically, the source node is used to generate a query message according to the context of the collective communication system 100, and the query message is used to request to query the in-network computing capabilities of the switch network 102. The switch 1020 is used to forward the query message transmitted from the source node to the destination node. The destination node is used to generate a notification message according to the query message, and the notification message carries the in-network computing capabilities of the switch network. The switch 1020 is further used to forward the notification message transmitted from the destination node to the source node.

[0339] In some possible implementation manners, the switch 1020 is specifically used for:

[0340] When receiving a query message, add the in-network computing capability of the switch 1020 to the query message;

[0341] Forward the query message with the in-network computing capability of the switch 1020 added to the destination node;

[0342] Correspondingly, the destination node is specifically used for:

[0343] Obtain the in-network computing capability of the switch network 102 according to the in-network computing capability of the switch 1020 in the query message forwarded by the switch 1020;

[0344] Generate a notification message according to the in-network computing capability of the switch network 102.

[0345] In some possible implementation manners, the in-network computing capability of the switch 1020 includes the set operation types and / or data types supported by the switch 1020.

[0346] In some possible implementation manners, the in-network computing capability of the switch 1020 further includes the size of the remaining available resources for in-network computing of the switch 1020.

[0347] In some possible implementation manners, the switch 1020 is further used for:

[0348] Establish an entry according to the hop count of the query message, and the entry is used for the switch 1020 to perform computing offloading on service messages.

[0349] In some possible implementation manners, when the switch 1020 is directly connected to the source node and the destination node, the switch 1020 is specifically used for:

[0350] Receive the query message sent by the source node and forward the query message to the destination node;

[0351] Receive the notification message sent by the destination node and forward the notification message to the source node.

[0352] In some possible implementation manners, the switch network 102 further includes a second switch and / or a third switch. The second switch is used to connect the switch 1020 and the source node, and the third switch is used to connect the switch 1020 and the destination node;

[0353] The switch 1020 is specifically used for:

[0354] Receive the query message sent by the source node and forward the query message to the third switch;

[0355] Receive the notification message forwarded by the third switch and forward the notification message to the source node; or,

[0356] Receive the query message forwarded by the second switch and forward the query message to the destination node;

[0357] Receive the notification message sent by the destination node and forward the notification message to the second switch; or,

[0358] Receive the query message forwarded by the second switch and forward the query message to the third switch;

[0359] Receive the notification message forwarded by the third switch and forward the notification message to the second switch.

[0360] In some possible implementation manners, the switch network 102 includes a single-layer switch, the switch 1020 is a single-layer switch, and the switch 1020 is specifically configured to:

[0361] Forward the notification message to the source node.

[0362] In some possible implementation manners, the switch network 102 includes an upper-layer switch and a lower-layer switch, and the switch 1020 is the lower-layer switch;

[0363] The switch 1020 is specifically configured to:

[0364] Determine a target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing of the upper-layer switch;

[0365] Add the size of the remaining available resources for in-network computing of the target switch to the notification message;

[0366] Forward the notification message added with the size of the remaining available resources for in-network computing of the target switch to the source node.

[0367] In some possible implementation manners, the switch 1020 is further configured to:

[0368] Determine a target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing of the upper-layer switch by using a load balancing strategy.

[0369] In some possible implementation manners, the switch 1020 is further configured to:

[0370] Send a switch query message to the upper-layer switch, where the switch query message is used to query the size of the remaining available resources for in-network computing of the upper-layer switch;

[0371] Receive the switch notification message sent by the upper-layer switch, where the switch notification message is used to notify switch 1020 of the size of the remaining available resources for in-network computing of the upper-layer switch.

[0372] In some possible implementation manners, the switch network 102 includes an upper-layer switch and a lower-layer switch, and switch 1020 is the upper-layer switch;

[0373] Switch 1020 is further configured to:

[0374] Receive the switch query message sent by the lower-layer switch, where the switch query message is used to query the size of the remaining available resources for in-network computing of switch 1020;

[0375] Send a switch notification message to the lower-layer switch, where the switch notification message is used to notify the lower-layer switch of the size of the remaining available resources for in-network computing of switch 1020.

[0376] In some possible implementation manners, the context of the collective communication system includes the context of the application or the context of the communication domain.

[0377] In some possible implementation manners, the multiple computing nodes include a master node and at least one sub-node;

[0378] The source node is the sub-node and the destination node is the master node; or,

[0379] The source node is the master node and the destination node is the sub-node.

[0380] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation manner. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0381] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0382] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0383] The above is only the specific implementation manner of the present application. Those skilled in the art of this technology can think of changes or substitutions according to the specific implementation manner provided by the present application, and all should be covered within the protection scope of the present application.

Claims

1. A method for processing control messages in a collective communication system, characterized in that, The described collective communication system includes a switch network and multiple computing nodes. The switch network includes a first switch. The method includes: The first switch forwards a query message transmitted from a source node to a destination node. The query message is used to request a query of the on-network computing capabilities of the switch network. The query message is generated by the source node by multiplexing the context of the collective communication system. The source node and the destination node are different nodes among the multiple computing nodes. The first switch forwards a notification message transmitted from the destination node to the source node. The notification message carries the on-network computing capabilities of the switch network. The on-network computing capabilities of the switch network include the on-network computing capabilities of the switches through which the query message passes.

2. The method according to claim 1, wherein The first switch forwarding a query message transmitted from a source node to a destination node includes: When the first switch receives a query message, it adds the on-network computing capabilities of the first switch to the query message. The first switch forwards the query message with the on-network computing capabilities of the first switch added.

3. The method according to claim 2, wherein The on-network computing capabilities of the first switch include the types of collective operations and / or data types supported by the first switch.

4. The method according to claim 2 or 3, characterized in that, The on-network computing capabilities of the first switch include the size of the remaining available resources for on-network computing of the first switch.

5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The first switch establishes an entry based on the hop count of the query message. The entry is used by the first switch for computing offloading of service messages.

6. The method according to any one of claims 1 to 3, characterized in that The first switch is directly connected to the source node and the destination node. The first switch forwarding a query message transmitted from a source node to a destination node includes: The first switch receives the query message sent by the source node and forwards the query message to the destination node. The first switch forwarding a notification message transmitted from the destination node to the source node includes: The first switch receives the notification message sent by the destination node and forwards the notification message to the source node.

7. The method according to any one of claims 1 to 3, characterized in that The switch network further includes a second switch and / or a third switch. The second switch is used to connect the first switch and the source node, and the third switch is used to connect the first switch and the destination node. The first switch forwarding a query message transmitted from a source node to a destination node includes: The first switch receives the query message sent by the source node and forwards the query message to the third switch; or, The first switch receives the query message forwarded by the second switch and forwards the query message to the destination node; or, The first switch receives the query message forwarded by the second switch and forwards the query message to the third switch. The first switch forwarding a notification message transmitted from the destination node to the source node includes: The first switch receives the notification message forwarded by the third switch and forwards the notification message to the source node; or, The first switch receives the notification message sent by the destination node and forwards the notification message to the second switch; or, The first switch receives the notification message forwarded by the third switch and forwards the notification message to the second switch.

8. The method according to any one of claims 1 to 3, characterized in that The switch network includes a single-layer switch. The first switch is the single-layer switch. The first switch transmits the notification message from the destination node to the source node, including: The first switch forwards the notification message to the source node.

9. The method according to any one of claims 1 to 3, characterized in that The switch network includes an upper-layer switch and a lower-layer switch. The first switch is the lower-layer switch; The first switch transmits the notification message from the destination node to the source node, including: The first switch determines a target switch from the upper-layer switches according to the size of the remaining available resources calculated online by the upper-layer switch; The first switch adds the size of the remaining available resources calculated online by the target switch to the notification message; The first switch forwards the notification message added with the size of the remaining available resources calculated online by the target switch to the source node.

10. The method according to claim 9, wherein The first switch determines a target switch from the upper-layer switches according to the size of the remaining available resources calculated online by the upper-layer switch, including: The first switch determines a target switch from the upper-layer switches according to the size of the remaining available resources calculated online by the upper-layer switch by using a load balancing strategy.

11. The method according to claim 9, wherein, The method further includes: The first switch sends a switch query message to the upper-layer switch. The switch query message is used to query the size of the remaining available resources calculated online by the upper-layer switch; The first switch receives a switch notification message sent by the upper-layer switch. The switch notification message is used to notify the first switch of the size of the remaining available resources calculated online by the upper-layer switch.

12. The method according to any one of claims 1 to 3, characterized in that, The switch network includes an upper-layer switch and a lower-layer switch. The first switch is the upper-layer switch; The method further includes: The first switch receives a switch query message sent by the lower-layer switch. The switch query message is used to query the size of the remaining available resources calculated online by the first switch; The first switch sends a switch notification message to the lower-layer switch. The switch notification message is used to notify the lower-layer switch of the size of the remaining available resources calculated online by the first switch.

13. The method according to any one of claims 1 to 3, characterized in that, The context of the collective communication system includes the context of the application or the context of the communication domain.

14. The method according to any one of claims 1 to 3, characterized in that, The multiple computing nodes include a master node and at least one slave node; The source node is the slave node and the destination node is the master node; or, The source node is the master node and the destination node is the slave node.

15. A method for processing control messages in a collective communication system, characterized in that, The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The multiple computing nodes include a first computing node and a second computing node. The method includes: The first computing node receives query messages forwarded by one or more switches in the switch network. The query messages are used to request a query of the in-network computing capabilities of the switch network. The query messages are generated by the second computing node by multiplexing the context of the collective communication system; The first computing node generates a notification message according to the query messages. The notification message carries the in-network computing capabilities of the switch network. The in-network computing capabilities of the switch network include the in-network computing capabilities of the switches through which the query messages pass; The first computing node sends the notification message to the second computing node through the one or more switches.

16. The method according to claim 15, wherein The in-network computing capabilities of the switch network carried by the notification message are obtained from the query messages forwarded by the one or more switches.

17. The method according to claim 15 or 16, characterized in that, The query messages forwarded by the switches include the in-network computing capabilities of the switches added by the switches. The method further includes: The first computing node obtains the in-network computing capabilities of the switch network according to the in-network computing capabilities of the one or more switches in the query messages forwarded by the one or more switches.

18. The method according to claim 15 or 16, characterized in that, The first computing node is a master node or a slave node.

19. A method for processing control messages in a collective communication system, characterized in that The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The multiple computing nodes include a first computing node and a second computing node. The method includes: The second computing node generates query messages according to the context of the collective communication system. The query messages are used to request a query of the in-network computing capabilities of the switch network; The second computing node sends the query messages to the first computing node through one or more switches in the switch network; The second computing node receives the notification message forwarded by the first computing node through the one or more switches. The notification message carries the in-network computing capabilities of the switch network.

20. The method according to claim 19, wherein The in-network computing capabilities of the switch network are obtained by the first computing node according to the in-network computing capabilities of the one or more switches in the query messages forwarded by the one or more switches.

21. The method according to claim 19 or 20, characterized in that, The second computing node is a master node or a slave node.

22. A control message processing device in a collective communication system, characterized in that, The collective communication system includes a switch network and multiple computing nodes. The switch network includes a first switch. The device includes: A communication module, configured to forward query messages transmitted from a source node to a destination node. The query messages are used to request a query of the in-network computing capabilities of the switch network. The query messages are generated by the source node by multiplexing the context of the collective communication system. The source node and the destination node are different nodes among the multiple computing nodes; The communication module is further configured to forward a notification message transmitted from the destination node to the source node. The notification message carries the in-network computing capabilities of the switch network. The in-network computing capabilities of the switch network include the in-network computing capabilities of the switches through which the query messages pass.

23. The device according to claim 22, characterized in that, The device further includes: A processing module, configured to add the in-network computing capabilities of the first switch to the query messages when the query messages are received; The communication module is specifically configured to: Forward a query message added with the in-network computing capacity of the first switch.

24. The device according to claim 23, characterized in that, The in-network computing capacity of the first switch includes the set operation types and / or data types supported by the first switch.

25. The device according to claim 23 or 24, characterized in that, The in-network computing capacity of the first switch includes the size of the remaining available resources for in-network computing of the first switch.

26. The device according to any one of claims 22 to 24, characterized in that, The device further includes: A processing module, configured to establish an entry according to the hop count of the query message, where the entry is used for the first switch to perform computing offloading on service messages.

27. The device according to any one of claims 22 to 24, characterized in that, The first switch is directly connected to the source node and the destination node; The communication module is specifically configured to: Receive a query message sent by the source node, and forward the query message to the destination node; Receive a notification message sent by the destination node, and forward the notification message to the source node.

28. The device according to any one of claims 22 to 24, characterized in that, The switch network further includes a second switch and / or a third switch. The second switch is used to connect the first switch and the source node, and the third switch is used to connect the first switch and the destination node; The communication module is specifically configured to: Receive a query message sent by the source node, and forward the query message to the third switch; Receive a notification message forwarded by the third switch, and forward the notification message to the source node; or, Receive a query message forwarded by the second switch, and forward the query message to the destination node; Receive a notification message sent by the destination node, and forward the notification message to the second switch; Or, Receive a query message forwarded by the second switch, and forward the query message to the third switch; Receive a notification message forwarded by the third switch, and forward the notification message to the second switch.

29. The device according to any one of claims 22 to 24, characterized in that, The switch network includes a single-layer switch, and the first switch is the single-layer switch; The communication module is specifically configured to: Forward the notification message to the source node.

30. The device according to any one of claims 22 to 24, characterized in that The switch network includes an upper-layer switch and a lower-layer switch, and the first switch is the lower-layer switch; The device further includes: A processing module, configured to determine a target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing of the upper-layer switch, and add the size of the remaining available resources for in-network computing of the target switch to the notification message. The communication module is specifically configured to: Forward the notification message added with the size of the remaining available resources for in-network computing of the target switch to the source node.

31. The device according to claim 30, wherein, The processing module is specifically configured to: Determine a target switch from the upper-layer switches according to the size of the remaining available resources for in-network computing of the upper-layer switch by using a load balancing strategy.

32. The device according to claim 30, wherein, The communication module is further configured to: Send a switch query message to the upper-layer switch, where the switch query message is used to query the size of the remaining available resources for in-network computing of the upper-layer switch; Receive a switch notification message sent by the upper-layer switch, where the switch notification message is used to notify the size of the remaining available resources for in-network computing of the upper-layer switch.

33. The device according to any one of claims 22 to 24, characterized in that, The switch network includes an upper-layer switch and a lower-layer switch, and the first switch is the upper-layer switch; The communication module is further configured to: Receive a switch query message sent by the lower-layer switch, where the switch query message is used to query the size of the remaining available resources for in-network computing of the first switch; Send a switch notification message to the lower-layer switch, where the switch notification message is used to notify the lower-layer switch of the size of the remaining available resources for in-network computing of the first switch.

34. A control message processing device in a collective communication system, characterized in that, The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The multiple computing nodes include a first computing node and a second computing node. The device includes: A communication module, configured to receive query messages forwarded by one or more switches in the switch network, where the query messages are used to request a query of the in-network computing ability of the switch network, and the query messages are generated by the second computing node by multiplexing the context of the collective communication system; A generation module, configured to generate a notification message according to the query messages, where the notification message carries the in-network computing ability of the switch network, and the in-network computing ability of the switch network includes the in-network computing ability of the switches through which the query messages pass; The communication module is further configured to send the notification message to the second computing node through the one or more switches.

35. The device according to claim 34, characterized in that, The query messages forwarded by the switch include the in-network computing ability of the switch added by the switch; Specifically, the generation module is configured to: Obtain the in-network computing ability of the switch network according to the in-network computing ability of the one or more switches in the query messages forwarded by the one or more switches; Generate a notification message according to the in-network computing ability of the switch network.

36. The device according to claim 34 or 35, characterized in that, The device is deployed on the first computing node, and the first computing node is a master node or a slave node.

37. A control message processing device in a collective communication system, characterized in that, The collective communication system includes a switch network and multiple computing nodes. The switch network includes at least one switch. The multiple computing nodes include a first computing node and a second computing node. The device includes: A generation module, configured to multiplex the context of the collective communication system to generate a query message, where the query message is used to request a query of the in-network computing ability of the switch network; A communication module, configured to send the query message to the first computing node through one or more switches in the switch network; The communication module is further configured to receive a notification message forwarded by the first computing node through the one or more switches, where the notification message carries the in-network computing ability of the switch network, and the in-network computing ability of the switch network includes the in-network computing ability of the switches through which the query message passes.

38. The device according to claim 37, characterized in that, The device is deployed on the second computing node, and the second computing node is a master node or a slave node.

39. A switch, characterized in that, The switch includes a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the switch executes the method according to any one of claims 1 to 14.

40. A computing node, characterized in that, The computing node includes a processor and a memory; The processor is configured to execute the instructions stored in the memory, so that the computing node executes the method according to any one of claims 15 to 18.

41. A computing node, characterized in that, The computing node includes a processor and a memory; The processor is configured to execute the instructions stored in the memory, so that the computing node executes the method according to any one of claims 19 to 21.

42. A collective communication system, characterized in that, The collective communication system includes a switch network and a plurality of computing nodes. The switch network includes a first switch, and the plurality of computing nodes include a first computing node and a second computing node; The second computing node is configured to multiplex a context generation query message for the collective communication system. The query message is used to request a query of the in-network computing capability of the switch network; The first switch is configured to forward the query message transmitted by the second computing node to the first computing node; The first computing node is configured to generate a notification message according to the query message. The notification message carries the in-network computing capability of the switch network. The in-network computing capability of the switch network includes the in-network computing capability of the switches through which the query message passes; The first switch is further configured to forward the notification message transmitted by the first computing node to the second computing node.

Citation Information

Patent Citations

  • Method for switching optimizing link of RRPP loop, system and network node

    CN101465782A

  • Multi-service adaptable routing protocol for wireless sensor networks

    US20110149844A1