Set communication processing method, device, computer device and storage medium

By updating the arrangement order of nodes in a distributed system, generating an orchestration node list, and performing collective communication according to the list, the problem of low collective communication efficiency in a distributed system is solved, and the effect of reducing communication delay and improving efficiency is achieved.

CN118869567BActive Publication Date: 2025-05-30TENCENT CLOUD COMPUTING (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410218312.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2025-05-30
Estimated Expiration
2044-02-27

AI Technical Summary

Technical Problem

In a distributed system, when each node based on distributed settings performs collective communication, the collective operation efficiency is low, resulting in an increase in communication delay.

Method used

By obtaining node orchestration requests and topology information, the arrangement order of nodes is updated step by step, a list of orchestration nodes is generated, and a collection communication is carried out according to the list.

Benefits of technology

Cross-node communication is reduced, and communication delay of the collection operation in the collection communication is reduced, thereby improving the efficiency of the collection operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118869567B_ABST
    Figure CN118869567B_ABST
Patent Text Reader

Abstract

The present application relates to a collective communication processing method, apparatus, computer device, storage medium, and computer program product. The method involves network scheduling and includes: obtaining a node orchestration request for a target collective operation in collective communication; the original node list carried in the node orchestration request includes each server node for the target collective operation; obtaining the topology information of each server node; the topology information is used to describe the switching nodes connected by the server nodes in each layer of the node communication network; each server node communicates with each other based on the switching nodes connected in each layer; based on each topology information, according to the layers in the node communication network, the arrangement order of each server node is gradually orchestrated and updated to obtain an orchestrated node list; according to the arrangement order of each server node in the orchestrated node list, collective communication is performed for the target collective operation through each server node. Using this method can improve the efficiency of collective operations in collective communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a collective communication processing method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] With the development of computer technologies, a distributed system composed of multiple independent nodes can greatly improve computing efficiency and save overall computing time. For example, for large-scale artificial intelligence models, efficient training can be performed based on a high-performance computing cluster composed of distributed servers.

[0003] In a distributed system, through various collective operations in collective communication, data can be efficiently exchanged, states can be synchronized, and behaviors can be coordinated among various nodes, so as to achieve collaborative work and global computing. However, currently, when performing collective communication among nodes based on a distributed setting, the efficiency of collective operations among various nodes is relatively low. Summary of the Invention

[0004] Based on this, it is necessary to provide a collective communication processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the efficiency of collective operations in collective communication for the above technical problems.

[0005] In a first aspect, the present application provides a collective communication processing method, including:

[0006] Obtaining a node orchestration request for a target collective operation in collective communication; the node orchestration request carries an original node list, and the original node list includes each server node for the target collective operation;

[0007] Obtaining the topology information of each server node; the topology information is used to describe the switching nodes connected by the server nodes at each level in the node communication network; each server node communicates based on the switching nodes connected at each level;

[0008] Based on each topology information, the arrangement order of each server node is gradually orchestrated and updated according to the levels in the node communication network to obtain an orchestrated node list;

[0009] Performing collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestrated node list.

[0010] In a second aspect, the present application further provides a collective communication processing apparatus, including:

[0011] A node orchestration request acquisition module, configured to acquire a node orchestration request for a target set operation in collective communication; the node orchestration request carries an original node list, and the original node list includes each server node for the target set operation.

[0012] A topology information acquisition module, configured to acquire the topology information of each server node; the topology information is used to describe the switching nodes connected by the server nodes at each level of the node communication network; each server node communicates based on the switching nodes connected at each level.

[0013] A node orchestration update module, configured to, based on each topology information, and in accordance with the levels in the node communication network, gradually arrange and update the arrangement order of each server node to obtain an arranged node list.

[0014] A collective operation processing module, configured to perform collective communication for the target collective operation through each server node according to the arrangement order of each server node in the arranged node list.

[0015] In a third aspect, the present application further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above collective communication processing method are implemented.

[0016] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above collective communication processing method are implemented.

[0017] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above collective communication processing method are implemented.

[0018] For each server node in the original node list carried by the node orchestration request for the target collective operation, the above collective communication processing method, device, computer device, storage medium, and computer program product arrange and update the arrangement order of each server node level by level based on the topology information of each server node in accordance with the levels in the node communication network to obtain an arranged node list, and perform collective communication for the target collective operation through each server node according to the arrangement order of each server node in the arranged node list. By arranging and updating the arrangement order of each server node in the original node list level by level according to the levels in the node communication network, and performing collective communication for the target collective operation based on the arrangement order of each server node after the arrangement update, cross-node communication can be reduced, the communication delay of collective operations in collective communication can be reduced, and thus the efficiency of collective operations can be improved.

[0019] In a sixth aspect, the present application provides a collective communication processing method, including:

[0020] When a target set operation in collective communication is triggered, determine an original node list; the original node list includes each server node for the target set operation.

[0021] Send a node orchestration request generated based on the original node list to the server; the node orchestration request is used to instruct the server to return an orchestrated node list; the orchestrated node list is obtained by orchestrating and updating the arrangement order of each server node in the original node list.

[0022] Receive the orchestrated node list, and perform collective communication for the target set operation through each server node according to the arrangement order of each server node in the orchestrated node list.

[0023] In a seventh aspect, the present application further provides a collective communication processing apparatus, including:

[0024] An original node list acquisition module, configured to determine an original node list when a target set operation in collective communication is triggered; the original node list includes each server node for the target set operation.

[0025] A node orchestration request module, configured to send a node orchestration request generated based on the original node list to the server; the node orchestration request is used to instruct the server to return an orchestrated node list; the orchestrated node list is obtained by orchestrating and updating the arrangement order of each server node in the original node list level by level according to the topology information of each server node and in accordance with the levels in the node communication network.

[0026] A set operation processing module, configured to receive the orchestrated node list, and perform collective communication for the target set operation through each server node according to the arrangement order of each server node in the orchestrated node list.

[0027] In an eighth aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above collective communication processing method are implemented.

[0028] In a ninth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above collective communication processing method are implemented.

[0029] In a tenth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above collective communication processing method are implemented.

[0030] The above set communication processing method, device, computer device, storage medium, and computer program product, when triggering a target set operation in set communication, determine an original node list including each server node for the target set operation, send a node orchestration request generated according to the original node list to the server, receive an orchestrated node list obtained by updating the arrangement order of each server node in the original node list returned by the server, and perform set communication for the target set operation through each server node according to the arrangement order of each server node in the orchestrated node list. For each server node for the target set operation, obtain an orchestrated node list obtained by hierarchically updating the arrangement order of each server node according to the hierarchy in the node communication network, and perform set communication for the target set operation according to the arrangement order of each server node in the orchestrated node list, which can reduce cross-node communication and reduce the communication delay of set operations in set communication, thereby improving the efficiency of set operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0032] Figure 1 It is an application environment diagram of the set communication processing method in some embodiments;

[0033] Figure 2 It is a flowchart of the set communication processing method in some embodiments;

[0034] Figure 3 It is a schematic diagram of data transmission by server nodes based on switching nodes in some embodiments;

[0035] Figure 4 It is a flowchart of the orchestration update process in some embodiments;

[0036] Figure 5 It is a flowchart of the set communication processing method in some embodiments;

[0037] Figure 6 It is a data center network topology diagram in some embodiments;

[0038] Figure 7 It is a data center network topology diagram in some embodiments;

[0039] Figure 8 For Figure 7 the logical traffic loop diagram in the illustrated embodiment;

[0040] Figure 9 For Figure 7 the network traffic schematic diagram in the illustrated embodiment;

[0041] Figure 10 For Figure 7 the logical traffic loop diagram without optimization in the illustrated embodiment;

[0042] Figure 11 For Figure 7 the network traffic schematic diagram without optimization in the illustrated embodiment;

[0043] Figure 12 the network topology diagram of a multi-tenant data center in some embodiments;

[0044] Figure 13 For Figure 12 the logical traffic loop diagram in the illustrated embodiment;

[0045] Figure 14 For Figure 12 the network traffic schematic diagram in the illustrated embodiment;

[0046] Figure 15 For Figure 12 the logical traffic loop diagram without optimization in the illustrated embodiment;

[0047] Figure 16 For Figure 12 the network traffic schematic diagram without optimization in the illustrated embodiment;

[0048] Figure 17 the relationship diagram between the number of upstream flows of the access layer switch and the probability of congestion in some embodiments;

[0049] Figure 18 the system architecture diagram of the collective communication processing method in some embodiments;

[0050] Figure 19 the process schematic diagram of the collective communication processing method in some embodiments;

[0051] Figure 20 the topology clustering schematic diagram in some embodiments;

[0052] Figure 21 the topology clustering result schematic diagram in some embodiments;

[0053] Figure 22 For Figure 21 the logical traffic loop diagram in the illustrated embodiment;

[0054] Figure 23 the net throughput comparison diagram of whether the collective communication processing is optimized or not in some embodiments;

[0055] Figure 24Structural block diagram of a collective communication processing device in some embodiments;

[0056] Figure 25 Structural block diagram of a collective communication processing device in some embodiments;

[0057] Figure 26 Internal structure diagram of a computer device in some embodiments. Detailed implementation manners

[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] The collective communication processing method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the computer devices in this application environment include servers such as a first server 102 and a second server 104. The first server 102 communicates with the second server 104 through a network. The data storage system can store the data that the servers to which they are connected need to process. The data storage system can be set up separately, integrated on the servers to which they are connected, or placed on the cloud or other servers. When triggering a target collective operation in collective communication, the first server 102 can send a node orchestration request for the target collective operation to the second server 104. The node orchestration request carries an original node list, and the original node list includes each server node for the target collective operation, and the first server 102 is determined from each server node. After receiving the node orchestration request, for each server node in the original node list carried by the node orchestration request for the target collective operation, the second server 104 updates the arrangement order of each server node level by level according to the topology information of each server node in the node communication network to obtain an orchestrated node list. The second server 104 can return the orchestrated node list to the first server 102, so that the first server 102 performs collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestrated node list.

[0060] In addition, when triggering a target collective operation in collective communication, the first server 102 determines an original node list including each server node for the target collective operation. The first server 102 sends a node orchestration request generated based on the original node list to the second server 104, and receives an orchestrated node list obtained by updating the arrangement order of each server node in the original node list returned by the second server 104. The first server 102 performs collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestrated node list. In some embodiments, the collective communication processing method may also be implemented solely by the first server 102, that is, the first server 102 directly performs an arrangement update process on the arrangement order of each of the server nodes, and performs collective communication based on the obtained orchestrated node list.

[0061] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. In addition, at least one of the first server 102 and the second server 104 in this application environment can be replaced by a terminal. The terminal can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc.

[0062] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed and applied Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system backing, which can only be achieved through cloud computing.

[0063] Cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources through the network in a on-demand and easily scalable manner; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services through the network in a on-demand and easily scalable manner. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balance. With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from previous parallel and distributed computing, the emergence of cloud computing will drive a revolutionary change in the entire Internet model and enterprise management model conceptually.

[0064] In an exemplary embodiment, as Figure 2 shown, a method for collective communication processing is provided. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the application of this method to Figure 1 the second server in as an example for illustration, it includes the following steps 202 to step 208. Among them:

[0065] Step 202, obtain a node orchestration request for a target set operation in collective communication; the node orchestration request carries an original node list, and the original node list includes each server node for the target set operation.

[0066] Among them, collective communications is a global communication operation that involves all processes in a process group and can handle a large amount of inter-process communication, synchronization, and computation. In a distributed system, collective communications is particularly crucial because a large amount of communication is often required between individual nodes. For example, in the distributed training process of deep learning, collective communications can be used to transmit a large amount of data such as network model weight parameters and a large number of temporary variables generated during the training process to improve the model training efficiency. The basic operations of collective communications include send, receive, copy, barrier synchronization among processes within a group, and inter-node process synchronization (signal + wait), etc. These basic operations can be combined to obtain various collective operations. Collective operations refer to global communication operations jointly completed by a group of processes, which allow all processes in the process group to participate in communication and jointly operate on data. Collective operations can specifically include, but are not limited to, broadcast, gather, allgather, scatter, reduce, allreduce, reducerscatter, and all-to-all, etc. The target collective operation refers to the collective operation that needs to be executed currently, such as an allreduce operation.

[0067] The node orchestration request is used to request an update to the arrangement order of each server node for the target collective operation. The server nodes are the nodes selected to implement the target collective operation, that is, the target collective operation is implemented among the selected server nodes. For example, an allreduce operation can be performed among the selected server nodes. The node orchestration request carries the original node list, and the original node list includes each server node selected to implement the target collective operation.

[0068] Specifically, the second server can obtain the node orchestration request for the target collective operation in collective communications. Specifically, it can receive the node orchestration request sent by the first server through the network. The first server can be a node among each server node for the target collective operation. As the server for updating the orchestration of each server node for the target collective operation, the second server can parse the original node list from the node orchestration request. The original node list includes each server node for the target collective operation. In a specific application, the original node list can record the node identifiers of each server node, such as the node name, node number, and other identification information of the server node.

[0069] Step 204: Obtain the topology information of each server node; the topology information is used to describe the switching nodes connected by the server node in each layer of the node communication network; each server node communicates based on the switching nodes connected in each layer.

[0070] Among them, the topology information describes the switching nodes connected by the server node in each layer of the node communication network, so as to perform communication at the corresponding layer through the connected switching nodes. The server node can be directly communicatively connected to the switching node, or indirectly communicatively connected through the switching nodes in other layers. For example, the topology information may include the node identifiers of the switching nodes connected by the server node in each layer of the node communication network. The node communication network may include multiple switching nodes for forwarding data packets between servers, and the switching nodes can be implemented based on devices such as terminals, servers, and switches. Each switching node in the node communication network can be constructed based on a multi-layer structure, and each layer may include at least one switching node, and each server node communicates based on the switching nodes connected in each layer. The data routing method when each server node communicates in the node communication network can be recorded, that is, data is transmitted through the corresponding switching nodes in the node communication network to achieve communication between different server nodes.

[0071] Exemplarily, for each server node in the original node list, the second server can obtain the topology information of each server node in the node communication network respectively. For example, the second server can query and determine the topology information of each server node in the node communication network based on the node identifiers of each server node. The node communication network is used to support communication between each server node, and may include multiple layers, so as to realize communication between server nodes through the switching nodes in multiple layers. The topology information may record the switching nodes connected by the server node in each layer. For example, the topology information may be {H0:1,1,1}, which describes that the server node H0 is connected to the switching node L1 in the first layer, the switching node P1 in the second layer, and the switching node C1 in the third layer in the node communication network. That is, in the node communication network, the server node H0 can communicate with other server nodes through the switching node L1 in the first layer, the switching node P1 in the second layer, and the switching node C1 in the third layer. Further, the switching nodes in the higher layer are used to realize cross-node communication between the server nodes connected by the switching nodes in the lower layer. For example, when realizing the communication across the switching node in the second layer between the server node connected by the switching node P1 in the second layer and the server node connected by the switching node P2 in the second layer, it needs to be realized based on the switching node in the third layer, such as can be realized based on the switching node C1 in the third layer.

[0072] In specific implementation, such as Figure 3As shown, the node communication network may include two layers, the first layer is the L layer, and the second layer is the S layer. The L layer includes the switching node L0 and the switching node L1. The switching node L0 is connected to the server nodes H0, H1, H2 and H3, and the switching node L1 is connected to the server nodes H4, H5, H6 and H7. The S layer can also be divided into S0 and S1 to support cross-node communication between the server nodes connected to L0 and L1. That is, the server nodes H0, H1, H2 and H3 are directly connected to the switching node L0 for communication, and are indirectly connected to the switching node S0 through the switching node L0. When the server node H0 and the server node H1 communicate, the server node H0 and the server node H1 are both connected to the switching node L0 of the L layer, then the communication between the server node H0 and the server node H1 can be directly based on the switching node L0 of the L layer, that is, the data packets between the server node H0 and the server node H1 can be forwarded through the switching node L0. When server node H0 and server node H7 communicate, they need to cross different switching nodes in the L layer, and they need to communicate across nodes through the switching nodes in the S layer. Specifically, the data packets between server node H0 and server node H7 can be forwarded through the switching nodes in the L layer, the switching nodes in the S layer, and the switching nodes in the L layer. Thus, through the switching nodes in the S layer, cross-node communication between server nodes connected to different switching nodes in the L layer is realized. Compared with the communication between server node H0 and server node H1 through the switching node L0, the communication between server node H0 and server node H7 needs to be realized through the switching nodes in the S layer across the switching nodes L0 and L1. The number of switching nodes involved in the communication process increases, which increases the communication delay.

[0073] Step 206 , based on each topology information, and according to the levels in the node communication network, the arrangement order of each server node is updated step by step to obtain an arrangement node list.

[0074] The arrangement node list is updated based on the original node list, and can be specifically re-arranged based on the arrangement order of each server node. When the arrangement order of each server node is updated step by step, it can be implemented from high to low or from low to high.

[0075] Optionally, the second server may reorder and update the arrangement order of each server node. Specifically, based on the respective topology information of each server, the reordering and updating are performed level by level in the hierarchy of the node communication network to reorder the arrangement order of each server node and obtain a reordered node list. The arrangement order of each server node in the reordered node list may be determined based on the hierarchical dimension of the reordering and updating performed level by level. By reordering and updating the arrangement order of each server node level by level in the hierarchy of the node communication network, it can be ensured that each server node in the reordered node list is arranged level by level in the hierarchy of the node communication network. Therefore, when performing collective communication on the target set operation according to the arrangement order of each server node, cross-node communication can be reduced, the communication delay of the collective operation in collective communication can be reduced, and thus the efficiency of the collective operation can be improved.

[0076] Step 208: Perform collective communication on the target set operation through each server node according to the arrangement order of each server node in the reordered node list.

[0077] Exemplarily, for the reordered node list, the arrangement order of each server node has been reordered and updated. The second server performs collective communication on the target set operation through each server node according to the arrangement order of each server node. For example, it may perform the target set operation in sequence by each server node according to the arrangement order of each server node in the reordered node list, such as performing a many-to-many reduction operation in sequence.

[0078] In the above collective communication processing method, for each server node in the original node list carried in the node reordering request for the target set operation, based on the topology information of each server node, the arrangement order of each server node is reordered and updated level by level in the hierarchy of the node communication network to obtain a reordered node list, and collective communication is performed on the target set operation through each server node according to the arrangement order of each server node in the reordered node list. By reordering and updating the arrangement order of each server node in the original node list level by level according to the hierarchy of the node communication network and performing collective communication on the target set operation based on the arrangement order of each server node after the reordering and updating, cross-node communication can be reduced, the communication delay of the collective operation in collective communication can be reduced, and thus the efficiency of the collective operation can be improved.

[0079] In an exemplary embodiment, as Figure 4 shown, the processing of the reordering and updating, that is, based on each topology information, the arrangement order of each server node is reordered and updated level by level in the hierarchy of the node communication network to obtain a reordered node list, includes:

[0080] Step 402: Based on each piece of topology information, hierarchically cluster each server node according to the levels in the node communication network to obtain the topology clustering results of each server node.

[0081] Among them, the switching nodes in the node communication network are divided into multiple levels. When the server nodes respectively connected by the switching nodes at the same level communicate with each other, they need to be realized through the switching nodes at the upper level. By hierarchically clustering each server node according to the levels of each switching node in the node communication network, specifically, the hierarchical clustering can be performed in the order from low to high or from high to low, and the topological structure relationship of each server node in the node communication network can be accurately determined. For example, the switching nodes required for communication between each server node can be determined. The topology clustering results may include the hierarchical clustering results of the server nodes at each level.

[0082] Exemplarily, the second server can hierarchically cluster each server node according to the levels in the node communication network. For example, when there are two levels in the node communication network, the second server can, based on the topology information of each server node, hierarchically cluster each server node in the order from the first level to the second level to obtain the topology clustering results of each server node. In specific implementation, the second server can, based on the topology information of each server node, hierarchically cluster each server node according to the first level, and divide the server nodes connected to the same interaction node at the first level into the same class, and multiple clustering results at the first level can be obtained. For each clustering result at the first level, the second server can, based on the topology information of each server node, hierarchically cluster each server node according to the second level, and divide the server nodes connected to the same interaction node at the second level into the same class to obtain multiple clustering results at the second level. After each clustering result at the first level is hierarchically clustered according to the second level, multiple clustering results at the second level can be obtained. After traversing each clustering result at the first level, the clustering results at the second level of each clustering result at the first level can be obtained, and the second server can obtain the topology clustering results of each server node according to the clustering results at the second level of each clustering result at the first level.

[0083] Step 404: Based on each topology clustering result, perform node search hierarchically according to the levels in the node communication network to obtain an orchestration node list; the arrangement order of the server nodes in the orchestration node list conforms to the hierarchical order of the node search.

[0084] Among them, the node search can be implemented by performing a hierarchical search in ascending order or descending order. The arrangement order of each server node in the orchestration node list matches the hierarchical order of the node search, that is, the arrangement order of each server node can be determined based on the hierarchical order of the node search. For example, in the orchestration node list, each server node can be arranged in sequence according to the hierarchical order of the node search.

[0085] Optionally, the second server can further perform a node search according to the topological clustering results of each server node, specifically performing a node search level by level in the node communication network. In a specific implementation, the hierarchical order of the node search can be the same as or different from the hierarchical order of the hierarchical clustering. For example, when performing hierarchical clustering, it can be implemented level by level in descending order in the node communication network, while the node search can be implemented level by level in ascending order in the node communication network. For example, when there are two levels in the node communication network, the second server can, based on each topological clustering result, perform a node search on each server node level by level in the order from the second level to the first level to obtain an orchestration node list. In a specific implementation, the second server can perform a node search based on each topological clustering result at the second level, that is, based on each topological clustering result, search for the topological clustering a1 connected to a certain switching node A1 at the second level and the topological clustering a2 connected to a certain switching node A2 at the second level. Each server node in the topological clustering a1 is connected to the switching node A1 at the second level, and each server node in the topological clustering a2 is connected to the switching node A2 at the second level; the second server respectively performs a node search at the first level in the topological clustering a1 and the topological clustering a2, that is, searches for the topological clustering b1 connected to the switching node B1 at the first level and the topological clustering b2 connected to the switching node B2 at the first level in the topological clustering a1, and searches for the topological clustering b3 connected to the switching node B3 at the first level and the topological clustering b4 connected to the switching node B4 at the first level in the topological clustering a2, so as to obtain {a1[(b1),(b2)], a2[(b3),(b4)]} through the search. Among them, (b1) and (b2) are obtained by further performing a node search at the first level based on a1, and (b3) and (b4) are obtained by further performing a node search at the first level based on a2, so as to obtain the orchestration node list as {(b1),(b2),(b3),(b4)}, that is, arranged in the order of (b1)-(b2)-(b3)-(b4).

[0086] In some embodiments, different node searches may yield different lists of orchestration nodes. For example, after searching for topology cluster a1 connected to a certain switching node A1 in the second tier and topology cluster a2 connected to a certain switching node A2 in the second tier based on each topology clustering result, further node search may be first performed on topology cluster a2 to obtain topology cluster b3 connected to switching node B3 in the first tier and topology cluster b4 connected to switching node B4 in the first tier, and then further node search is performed on topology cluster a1, so that the list of orchestration nodes can be obtained as {(b3), (b4), (b1), (b2)}, that is, arranged in the order of (b3)-(b4)-(b1)-(b2). In specific applications, when (b1), (b2), (b3), and (b4) include multiple server nodes, the multiple server nodes in each of (b1), (b2), (b3), and (b4) can be arbitrarily sorted. For example, when b1 includes a total of three server nodes b11, b12, and b13, within (b1) in the list of orchestration nodes, it can be (b11, b12, b13), (b11, b13, b12), (b12, b11, b13), (b13, b11, b12), (b12, b13, b11), or (b13, b12, b11).

[0087] In this embodiment, the second server hierarchically clusters each server node based on the topology information of each server node according to the levels in the node communication network, and performs node search step by step based on the topology clustering results obtained by hierarchical clustering, so as to obtain an orchestration node list in which the arrangement order of each server node conforms to the level order of node search, which can ensure that each server node in the orchestration node list is arranged step by step according to the levels in the node communication network. Therefore, when performing collective communication on the target set operation through each server node in the arrangement order of each server node, cross-node communication can be reduced, the communication delay of the collective operation in collective communication can be reduced, and thus the efficiency of the collective operation can be improved.

[0088] In an exemplary embodiment, based on each topology information, each server node is hierarchically clustered step by step according to the levels in the node communication network to obtain the topology clustering results of each server node, including: based on each topology information, each server node is hierarchically clustered step by step according to the level dimension from high to low in the node communication network to obtain the topology clustering results of each server node; in the node communication network, high-level switching nodes are used to support cross-node communication between server nodes connected to low-level switching nodes.

[0089] Among them, multiple levels are set in the node communication network. Each level includes at least one switching node. The switching nodes at higher levels are used to support cross-node communication between server nodes connected by the switching nodes at lower levels. That is, when node communication is performed between server nodes connected by different switching nodes in the same lower level, it is necessary to cross the switching nodes of this level, which is specifically implemented through the switching nodes at higher levels. For example, for switching nodes A and B in the same level, switching node A is connected to server node 1, and switching node B is connected to server node 2. When communication is performed between server node 1 and server node 2, since in this level, server node 1 and server node 2 are respectively connected to different switching nodes and need to cross the switching nodes in this level for communication, it can be implemented through switching node C in the upper level of this level. Then the specific data transmission route can be server node 1 - switching node A - switching node C - switching node B - server node 2. For hierarchical clustering step by step, it can be implemented according to the hierarchical dimension from high to low or from low to high.

[0090] Exemplarily, the second server can determine the hierarchical order of hierarchical clustering step by step. Specifically, based on the topological information of the server nodes, each server node can be hierarchically clustered step by step according to the hierarchical dimension from high to low to obtain the topological clustering results of each server node. Each topological clustering result can include the clustering results of the hierarchical clustering of the server nodes at each level. For example, when there are three levels in the node communication network, hierarchical clustering can be performed step by step according to the hierarchical dimension from the first level to the second level to the third level, or according to the hierarchical dimension from the third level to the second level to the first level.

[0091] In this embodiment, the second server hierarchically clusters each server node according to the hierarchical dimension from high to low, which can accurately determine the clustering results of each server node at each level, is conducive to accurately arranging each server node step by step according to the levels in the node communication network, and further reduces the communication between cross nodes, reduces the communication delay of the collective operation in the collective communication, and thus improves the efficiency of the collective operation.

[0092] In an exemplary embodiment, the levels in the node communication network include, from high to low, a core layer, an aggregation layer, and an access layer; based on each topological information, in accordance with the level dimension from high to low in the node communication network, each server node is hierarchically clustered step by step to obtain a topological clustering result of each server node, including: clustering each server node based on each topological information in accordance with the level dimension of the core layer to obtain a core layer clustering result of each server node; clustering each server node based on each core layer clustering result in accordance with the level dimension of the aggregation layer to obtain an aggregation layer clustering result of each server node; clustering each server node based on each aggregation layer clustering result in accordance with the level dimension of the access layer to obtain a topological clustering result of each server node.

[0093] Among them, the levels in the node communication network include, from high to low, a core layer (Core), an aggregation layer (Spine), and an access layer (Leaf). The core layer can be the third level, the aggregation layer is the second level, and the access layer is the first level. That is, the switching nodes of the core layer are used to support cross-node communication between server nodes connected by the switching nodes of the second level or the first level; the switching nodes of the aggregation layer are used to support cross-node communication between server nodes connected by the switching nodes of the first level. When hierarchically clustering step by step, it can be implemented in the level order of the core layer, the aggregation layer, and the access layer.

[0094] Exemplarily, for a node communication network including, from high to low, a core layer, an aggregation layer, and an access layer, the second server can cluster each server node based on each topological information in accordance with the level dimension of the core layer. Specifically, the server nodes connected to the same switching node in the core layer are divided into the same cluster to obtain a core layer clustering result of each server node. The second server clusters each server node based on each core layer clustering result in accordance with the level dimension of the aggregation layer. Specifically, the server nodes connected to the same switching node in the aggregation layer are divided into the same cluster to obtain an aggregation layer clustering result. The second server clusters each server node based on each aggregation layer clustering result in accordance with the level dimension of the access layer. Specifically, the server nodes connected to the same switching node in the access layer are divided into the same cluster to obtain a topological clustering result of each server node.

[0095] In this embodiment, for a node communication network including, from high to low, a core layer, an aggregation layer, and an access layer, the second server clusters each server node in sequence according to the level order of the core layer, the aggregation layer, and the access layer to obtain a topological clustering result reflecting the clustering results at each level, which is beneficial to accurately arranging each server node step by step according to the levels in the node communication network, thereby reducing cross-node communication and reducing the communication delay of set operations in collective communication, and thus improving the efficiency of set operations.

[0096] In an exemplary embodiment, based on each topological clustering result, node search is performed level by level in the node communication network to obtain an orchestration node list, including: based on each topological clustering result, performing node search for each server node level by level according to the hierarchical dimension from high to low in the node communication network to obtain an orchestration node list; in the node communication network, high-level switching nodes are used to support cross-node communication between server nodes connected by low-level switching nodes.

[0097] Among them, multiple levels are set in the node communication network, and each level includes at least one switching node. High-level switching nodes are used to support cross-node communication between server nodes connected by low-level switching nodes. That is, when node communication is performed between server nodes connected by different switching nodes in the same low level, the switching nodes of this level need to be crossed, and it is specifically implemented through high-level switching nodes.

[0098] Specifically, the second server can determine the hierarchical order of the step-by-step node search. Specifically, based on the topological clustering results of each server node, each server node is searched step by step according to the hierarchical dimension from high to low to obtain the orchestration node list of each server node. For example, when the node communication network includes three levels, the topological clustering results can be searched step by step according to the hierarchical dimension from the first level to the second level to the third level, or can be searched step by step according to the hierarchical dimension from the third level to the second level to the first level.

[0099] In this embodiment, the second server performs node search for each server node according to the hierarchical dimension from high to low, arranges each server node accurately level by level in the node communication network, thereby reducing cross-node communication, reducing the communication delay of collective operations in collective communication, and thus improving the efficiency of collective operations.

[0100] In an exemplary embodiment, the levels in the node communication network include the core layer, aggregation layer, and access layer from high to low; based on each topological clustering result, performing node search for each server node level by level according to the hierarchical dimension from high to low in the node communication network to obtain an orchestration node list, including: performing search according to the hierarchical dimension of the core layer based on each topological clustering result to obtain the core layer search result; performing search according to the hierarchical dimension of the aggregation layer based on the core layer search result to obtain the aggregation layer search result; performing search according to the hierarchical dimension of the access layer based on the aggregation layer search result to obtain the access layer search result; and obtaining the orchestration node list according to the core layer search result, aggregation layer search result, and access layer search result.

[0101] Among them, the levels in the node communication network include the core layer, the aggregation layer, and the access layer from high to low. Optionally, for a node communication network including the core layer, the aggregation layer, and the access layer from high to low, the second server may search based on each topological clustering result in the hierarchical dimension of the core layer. Specifically, for each switching node in the core layer, based on each topological clustering result, search for server nodes connected to the corresponding switching node to obtain the core layer search result. Among them, the core layer search result may record the server nodes connected to the switching nodes in the core layer; the server nodes connected to the switching nodes may be directly communicatively connected or indirectly communicatively connected through switching nodes in other levels. The second server searches based on the core layer search result in the hierarchical dimension of the aggregation layer. Specifically, for each switching node in the aggregation layer, based on each core layer search result, search for server nodes connected to the corresponding switching node to obtain the aggregation layer search result. The aggregation layer search result may include the server nodes connected to the switching nodes in the aggregation layer. The second server searches based on the aggregation layer search result in the hierarchical dimension of the access layer. Specifically, for each switching node in the access layer, based on each aggregation layer search result, search for server nodes connected to the corresponding switching node to obtain the access layer search result. The access layer search result may include the server nodes connected to the switching nodes in the access layer.

[0102] The second server can obtain an orchestration node list by integrating the core layer search results, aggregation layer search results, and access layer search results. Specifically, the second server can sort and combine each server node according to the core layer, aggregation layer, and access layer hierarchical search order, based on the core layer search results, aggregation layer search results, and access layer search results, to obtain an orchestration node list. For example, the core layer search results include c1, c2, and c3, which respectively correspond to the switching nodes C1, C2, and C3 in the core layer; for c1, c2, and c3, in the aggregation layer search results obtained by searching respectively according to the hierarchical dimension of the aggregation layer, s1 and s2 are obtained based on c1 search, s3 and s4 are obtained based on c2 search, s5 and s6 are obtained based on c3 search, and s1, s2, s3, s4, s5, and s6 respectively correspond to the switching nodes S1, S2, S3, S4, S5, and S6 in the aggregation layer; for s1, s2, s3, s4, s5, and s6, in the access layer search results obtained by searching respectively according to the hierarchical dimension of the access layer, l1 and l2 are obtained based on s1 search, l3 and l4 are obtained based on s2 search, l5 and l6 are obtained based on s3 search, l7 and l8 are obtained based on s4 search, l9 and l10 are obtained based on s5 search, l11 and l12 are obtained based on s6 search, and l1, l2, l3, l4, l5, l6, l7, l8, l9, l10, l11, and l12 respectively correspond to the switching nodes L1, L2, L3, L4, L5, L6, L7, L8, L9, L10, L11, and L12 in the access layer. The second server can sort each server node according to c1, c2, and c3, s1, s2, s3, s4, s5, and s6, as well as l12, l1, l2, l3, l4, l5, l6, l7, l8, l9, l10, l11, and l12, to obtain an orchestration node list in which the sorting order of each server node conforms to the hierarchical order of node search. For example, the orchestration node list can be {l1, l2, l3, l4, l5, l6, l7, l8, l9, l10, l11, l12}, or it can be {l12, l11, l10, l9, l8, l7, l6, l5, l4, l3, l2, l1}.

[0103] In this embodiment, for a node communication network including a core layer, an aggregation layer, and an access layer from high to low, the second server sequentially searches the topology clustering results according to the hierarchical order of the core layer, the aggregation layer, and the access layer to obtain an orchestration node list, which can arrange each server node level by level in the node communication network, thereby reducing communication between cross nodes and reducing the communication delay of set operations in collective communication, thus improving the efficiency of set operations.

[0104] In an exemplary embodiment, obtaining a node orchestration request for a target collective operation in collective communication includes: determining a collective communication library in collective communication; and obtaining a node orchestration request for a target collective operation in collective communication based on a communication interface with the collective communication library.

[0105] Among them, the Collective Communication Library is a library used to support communication between multiple processes in parallel computing. In a parallel computing environment, multiple processes may need to work together to complete a task, and the collective communication library provides efficient collective operations, enabling these processes to conveniently perform data exchange, synchronization, and other operations. Exemplarily, the second server can determine the collective communication library in collective communication. The collective communication library includes various collective operations, that is, corresponding collective operations are implemented based on the collective communication library. The collective communication library may include, but is not limited to, communication libraries such as MPI (Message Passing Interface), NCCL (NVIDIA Collective Communications Library), Libfabric (asynchronous communication library), or TBB (Intel Threading Building Blocks). The second server obtains a node orchestration request for a target collective operation through a communication interface with the collective communication library, and the second server can be implemented based on a plugin of the collective communication library through the communication interface with the collective communication library.

[0106] In this embodiment, the second server obtains a node orchestration request for a target collective operation through a communication interface with the collective communication library, so as to accurately and quickly obtain a node orchestration request including an original node list, which is beneficial to improving the collective communication efficiency for the target collective operation.

[0107] In an exemplary embodiment, obtaining topology information of each server node includes: determining a topology database associated with a node communication network; and querying the topology information of each server node from the topology database according to the original node list.

[0108] Among them, the topology database is used to record the hierarchical topology structure in the node communication network and the server nodes connected to each switching node. Specifically, the second server can determine the topology database associated with the node communication network, and query the topology information of each server node from the topology database according to the original node list carried in the node orchestration request. In a specific implementation, the original node list may include the node identifiers of each server node for the target set operation, and the second server can query the topology information of the identified server node in the topology database according to the node identifier. The topology information may include the node identifier of the server node and the node identifier of the connected switching node.

[0109] In this embodiment, the second server queries the topology information of each server node from the topology database associated with the node communication network according to the original node list, so that it can use the topology information of each server node to update the arrangement order of each server node step by step according to the hierarchy in the node communication network, and perform collective communication for the target set operation based on the arrangement order of each server node after the orchestration update, which can reduce the generation of cross-node communication, reduce the communication delay of the set operation in collective communication, and thus improve the efficiency of the set operation.

[0110] In an exemplary embodiment, obtaining a node orchestration request for a target set operation in collective communication includes: obtaining a node orchestration request from a target server node; the target server node is determined from each server node for the target set operation in collective communication.

[0111] Among them, the target server node is determined from each server node for the target set operation, and specifically, it can be the server node that initiates collective communication for the target set operation among each server node. Exemplarily, the second server determines the target server node and obtains a node orchestration request from the target server node. In a specific implementation, the node orchestration request can be actively sent by the target server node to the second server, and the second server can directly receive the node orchestration request sent by the target server node.

[0112] Further, performing collective communication for the target set operation through each server node according to the arrangement order of each server node in the orchestrated node list includes: sending the orchestrated node list to the target server node; the orchestrated node list is used to instruct the target server node to perform collective communication for the target set operation through each server node according to the arrangement order of each server node in the orchestrated node list.

[0113] Specifically, after the original node list carried in the node orchestration request sent by the target server node is orchestrated and updated, the second server may send the orchestrated node list to the target server node, so that the target server node performs collective communication for the target set operation according to the arrangement order of each server node in the orchestrated node list, that is, the target server node initiates the collective communication processing for the target set operation.

[0114] In this embodiment, the second server obtains the node orchestration request from the target server node and sends the orchestrated node list to the target server node, so that the target server node initiates collective communication for the target set operation based on the orchestrated node list, which can reduce cross-node communication and reduce the communication delay of the set operation in collective communication, thereby improving the efficiency of the set operation.

[0115] In an exemplary embodiment, the collective communication processing method further includes: when each server node for the target set operation satisfies the orchestration update trigger condition, performing the step of obtaining the node orchestration request for the target set operation in collective communication.

[0116] Among them, the orchestration update trigger condition is used to determine whether to trigger the orchestration update processing for the server node. For example, when the server node for the target set operation is updated, it is considered to satisfy the orchestration update trigger condition. For example, when there is a faulty node among the server nodes for the target set operation, or a new server node is added, it can be considered to satisfy the orchestration update trigger condition. Exemplarily, the second server may monitor each server node for the target set operation. Specifically, it may detect the state changes of each server node. When it detects that the orchestration update trigger condition is satisfied, such as when the second server monitors that there is a faulty node among the server nodes for the target set operation, or a new server node is added, the server performs the step of obtaining the node orchestration request for the target set operation in collective communication to timely orchestrate and update the arrangement order of each server node whose state has changed.

[0117] In this embodiment, when it is detected that the orchestration update trigger condition is satisfied, the second server performs the step of obtaining the node orchestration request for the target set operation in collective communication, which can timely orchestrate and update the arrangement order of each server node whose state has changed, thereby facilitating the improvement of the set operation efficiency.

[0118] In an exemplary embodiment, as Figure 5 shown, a collective communication processing method is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of the present application, this method is applied to Figure 1Taking the first server in as an example, the following steps 502 to 506 are included. Among them:

[0119] Step 502, when a target collective operation in collective communication is triggered, determine an original node list; the original node list includes each server node for the target collective operation.

[0120] Among them, the target collective operation refers to the collective operation that needs to be executed currently triggered in collective communication. The original node list includes each server node selected to implement the target collective operation. The server node is the node selected to implement the target collective operation, that is, the target collective operation is implemented among the selected server nodes.

[0121] Exemplarily, when a target collective operation in collective communication is triggered, indicating that the target collective operation needs to be executed, the first server can determine the corresponding original node list, and the original node list includes each server node for the target collective operation. The first server belongs to the nodes among each server node for the target collective operation.

[0122] Step 504, send a node orchestration request generated according to the original node list to the server; the node orchestration request is used to instruct the server to return an orchestrated node list; the orchestrated node list is obtained by orchestrating and updating the arrangement order of each server node in the original node list level by level according to the topology information of each server node in the node communication network.

[0123] Among them, the node orchestration request is generated based on the original node list, and the node orchestration request is used to request the arrangement order of each server node for the target collective operation to be orchestrated and updated. The topology information describes the switching nodes connected by the server nodes at each level in the node communication network, so as to perform corresponding level communication through the connected switching nodes. There can be multiple switching nodes in the node communication network for forwarding data packets between servers. The switching nodes in the node communication network can be constructed based on a multi-level structure. Each level can include at least one switching node, and the server nodes communicate with each other based on the switching nodes connected at each level. When orchestrating and updating the arrangement order of each server node level by level, it can be realized according to the hierarchical dimension from high to low or from low to high.

[0124] Optionally, the first server generates a node orchestration request according to the original node list and sends the node orchestration request to the server to instruct the server to orchestrate and update the arrangement order of each server node in the original node list level by level according to the topology information of each server node in the node communication network, and return the orchestrated node list obtained by the orchestration update.

[0125] Step 506: Receive the orchestration node list, and perform collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestration node list.

[0126] Specifically, for the orchestration node list, the arrangement order of each server node is re-arranged and updated, and the first server performs collective communication for the target set operation through each server node according to the arrangement order of each server node. For example, the first server can initiate the target set operation according to the arrangement order of each server node in the orchestration node list to control each server node to perform the target set operation in sequence, such as performing many-to-many protocol operations in sequence.

[0127] In the above-mentioned collective communication processing method, when the target collective operation in the collective communication is triggered, the original node list including each server node for the target collective operation is determined, the node orchestration request generated according to the original node list is sent to the server, and the orchestration node list returned by the server is received, which is obtained by arranging and updating the arrangement order of each server node in the original node list, and collective communication is performed for the target collective operation through each server node according to the arrangement order of each server node in the orchestration node list. For each server node for the target collective operation, an orchestration node list is obtained by arranging and updating the arrangement order of each server node step by step according to the hierarchy in the node communication network, and collective communication is performed for the target collective operation according to the arrangement order of each server node in the orchestration node list, which can reduce the generation of cross-node communication, reduce the communication delay of collective operation in collective communication, and thus improve the efficiency of collective operation.

[0128] The present application also provides an application scenario, and the application scenario applies the above-mentioned collective communication processing method. Specifically, the application of the collective communication processing method in the application scenario is as follows:

[0129] Modern data center networks are generally Clos architectures, that is, three-layer or two-layer switching nodes are set up, and the switching nodes can be switches. Figure 6 As shown, it includes three layers of switches, from the lowest layer to the highest layer, namely the access layer, aggregation layer and core layer, and a layer of server nodes, where the server nodes are specifically servers (Host). In the core layer, there are 8 switches C0-C7, the aggregation layer includes 16 switches S0-S15, the access layer includes 16 switches L0-L15, and the server layer includes 32 servers H0-H31. Among them, the switch is only responsible for forwarding data packets, and all business applications are deployed on the server. The switches at the access layer and the servers connected to it are collectively called a rack (Rack or block); and a group of interconnected aggregation layer and access layer switches and their downstream servers form a network module (Pod or module). For example, inFigure 6 In this figure, the access layer switch L0 and the servers H0 and H1 connected to it form a rack; and all the devices within the first dashed box form a Pod. It can be seen that for cross-Rack server communication, it is necessary to go up to the aggregation layer switch. For example, the communication between H0 and H2 needs to pass through one or more switches among S0 to S3. And for cross-Pod server communication, it is necessary to go up to the core layer switch. For example, the communication between H0 and H8 needs to pass through one or more switches among the core layer switches C0 to C7.

[0130] Currently, the mainstream large-scale artificial intelligence (AI) models are all trained using high-performance computing clusters (HPC, High Performance Computing) composed of distributed servers in the data center network. In such a computing cluster, multiple servers equipped with high-performance GPUs (Graphics Processing Unit) cooperate in training through high-bandwidth networks, such as RoCEv2 (RDMA over Converged Ethernet) network or IB (InfiniBand network). Among them, RDMA refers to Remote Direct Memory Access, that is, remote direct memory data access. In a typical training task, the master server coordinates a group of servers for training based on a certain training framework, such as TensorFlow or PyTorch. The load of cluster training includes computing tasks and communication tasks. The computing task refers to the mathematical calculation of the AI model by the GPU of the server; while the communication task mainly refers to synchronizing and coordinating the calculation results of each server, which is reflected as data transmission between servers. Usually, communication tasks are implemented based on collective communication operations in collective communication libraries (such as MPI, NCCL, libfabric, etc.), such as Allreduce (global reduction), Alltoall, Allgather, etc. Among them, NCCL (Nvidia Collective Communication Library) is the most widely used collective communication library in the current large model training industry, and Allreduce is the most basic and most used collective communication operation. In Allreduce, each node sends its own data, also receives data sent by other nodes in the group, and generates a summary result. This embodiment mainly focuses on the Allreduce operation and improves its throughput through data flow scheduling and dispatching.

[0131] The implementation of Allreduce generally uses the Parameter-Server (PS) architecture in traditional training. However, the completion time of the Allreduce operation in the PS architecture is proportional to the number of nodes, which means it is difficult to scale to large-scale networks. The peer-to-peer architecture emerged as a result, and this architecture can very well solve the scalability problem. In the peer-to-peer architecture, training mainly relies on distributed Allreduce collective operations to complete data synchronization. The implementation of distributed Allreduce operations includes Ring Allreduce (ring reduction operation), Tree Allreduce (tree reduction operation), Halving-Doubling Allreduce (halving-doubling reduction operation), etc. and their variants. Among them, Ring Allreduce is the most widely used Allreduce implementation, and NCCL is based on the peer-to-peer architecture and uses distributed Allreduce operations for parameter synchronization.

[0132] In the logical implementation of Ring Allreduce, the master server strings all the servers (usually a server has multiple GPUs) into a loop. Data is transmitted along the loop from one server to the next in turn until all servers have received the data from all other servers. The amount of data transmitted for each communication pair (such as H1->H2) is the same. In this loop, the completion time of the Allreduce communication operation is equal to the transmission time of the slowest communication pair. That is, a decrease in the rate of any communication pair will become the bottleneck of this collective communication and slow down the entire collective communication operation. In an actual network, such as Figure 7 In it, there are 8 servers H0~H7 participating in Ring Allreduce. Servers H0, H1, H2, and H3 are directly connected to the access layer switch L0, while servers H4, H5, H6, and H7 are directly connected to the access layer switch L1. When processing based on Ring Allreduce, the optimal logical traffic loop formed by them is as Figure 8 shown. In the optimal logical traffic loop of Ring Allreduce, the solid arrows are the traffic within the rack, and the dashed arrows are the traffic across the rack. Data is transmitted from the direction pointed by the arrow to the direction pointed to by the arrow. And the corresponding actual network traffic is as Figure 9As shown, it can be seen that in this ring, there are 6 intra-Rack communications, and the other 2 communications are inter-Rack communications. Intra-Rack communications only pass through one-hop switches and there is no risk of congestion (rate degradation); while inter-Rack communications may experience rate degradation due to sharing physical links with other data flows, and these other data flows may also belong to this Allreduce operation. There are two risks for inter-Rack traffic: one is the easy congestion with other traffic leading to rate reduction, and the other is that the latency is significantly higher than intra-Rack communications. Therefore, in model training, it is necessary to minimize inter-Rack traffic in an Allreduce operation. As Figure 8 and its corresponding network traffic Figure 9 is the optimal traffic scheduling scenario, where the inter-Rack traffic (dashed line) has been minimized, with only one set of forward (H3->H4) and one set of reverse (H7->H0).

[0133] In an actual large-scale production environment, especially in a multi-tenant HPC training cluster provided by cloud providers, there are characteristics of cloud environment service deployment, that is, the server nodes provided to tenants may not be arranged continuously in terms of physical location, and the server list provided to tenants is often in a disordered order. The main reasons are as follows:

[0134] Masking the physical location from tenants; in order to prevent tenants from making requests for server physical locations that are difficult to meet (such as requests within a Rack), and at the same time to maximize server resource utilization, the provider needs to present the effect that "all server resources have equivalent capabilities" to tenants.

[0135] The dynamics of server leasing are very high; since tenants can flexibly increase or decrease the number of leased servers at any time, the locations of idle / rentable servers are likely to be very scattered. In this way, if existing tenants subsequently want to incrementally purchase servers or new tenants purchase servers, the physical locations of the servers they obtain are likely to be discontinuous. Additionally, due to the order of purchase time, it may also result in a disordered order in the list even if the server locations are continuous.

[0136] Server failure replacement; Due to the high performance and power requirements of the HPC training scenario, the failure rate of these servers (including GPU failures, network card failures, and whole machine failures, etc.) is relatively high, and there are often scenarios where failed servers need to be replaced. Due to time urgency, server replacement is carried out using the redundant hot standby replacement method. That is, several servers are deliberately left unsold in the cluster. When a server fails, the redundant server is provided to the tenant for use, while the failed server remains in place, waiting for physical repair or replacement. After physical repair or replacement is completed, it will be sold again. In this way, for the tenant, the newly added server will be directly added to the existing server list (at the end). Therefore, the server list obtained by the tenant is often very chaotic in terms of physical distribution.

[0137] Existing collective communication libraries (such as NCCL) can query and sense the GPU interconnection topology within a node through the operating system and optimize it, but they do not have the ability to sense the topological order among multiple server nodes. The current situation is that tenants generally receive a list of server nodes from a cloud operator, pass this list to the collective communication library, and then the communication library arranges the nodes into a traffic topology ring according to the node order in the list for Ring Allreduce operations. However, as Figure 8 and Figure 9 show, it corresponds to the most ideal situation. In actual business scenarios, limited by the above cloud environment deployment capabilities, the node lists received by tenants are often out of order. In this case, the actual efficiency of collective operations may drop significantly (the typical value of the drop ratio is 50%). It should be noted that although Ring Allreduce is the most commonly used Allreduce implementation, other implementations of Allreduce are also applied, such as Tree Allreduce, Halving-Doubling Allreduce, etc. Similarly, these Allreduce implementations are also affected by the topological order of the servers, and this embodiment is equally applicable to these implementations.

[0138] From this, it can be seen that the performance of distributed Ring Allreduce depends very much on excellent topological affinity and network capabilities. Under the optimal topological arrangement, as Figure 8 shows, the machines under a single rack will only generate a set of upward traffic and a set of downward traffic. As Figure 9As shown in the figure, the 4 servers (H0~H3) of switch L0 in rack 1 only generate a set of upstream traffic (dashed line) passing through S1 to rack 2, and the 4 servers (H4~H7) of switch L1 in rack 2 also only generate a set of upstream traffic (dashed line) passing through S0 to rack 1. Note that the upstream traffic of rack 1 is the downstream traffic of rack 2, and vice versa. In this way, since the single-group traffic of each rack only originates from one server, it will not cause congestion in any upstream link of the access layer switch, and the network is naturally very smooth.

[0139] However, in actual training, if the topology received by the communication library is not the optimal topology, the formed traffic loop is often not the minimum loop, and there will be multiple groups of traffic spanning the aggregation layer and even the core layer switches in the loop. For example, Figure 10 As shown in the figure, in the traffic logic loop of the 8 servers H0-H7, as the worst-case scenario for a single tenant, the communication between all servers is traffic spanning multiple levels. As Figure 11 shown in the figure, in the traffic topology loop of the worst-case scenario for a single tenant, the communication between all servers spans the switches of the access layer and is implemented through the switches of the aggregation layer. Due to the defect of switch load balancing (the routing hash is random), these multiple groups of upstream traffic often cause a certain degree of congestion, resulting in a decrease in throughput between some node pairs. And the characteristics of the loop traffic determine that the overall performance of the system will be limited by the connection with the minimum throughput (bottleneck). Therefore, the congestion caused by non-optimal traffic topologies will seriously affect the overall performance of the HPC system. For example, if the arranged logical topology loop is Figure 10 as shown in the figure, that is, the loop is H0-H4-H1-H5-H2-H6-H3-H7-H0, then 4 groups of cross-rack upstream traffic (through the aggregation layer switches) and 4 groups of cross-rack downstream traffic will be generated in the network. The following will discuss in two cases.

[0140] If the network bandwidth is convergent, that is, the total downstream bandwidth of L0 is greater than the total upstream bandwidth, then these multiple groups of traffic will inevitably cause serious congestion, generating a bottleneck with very poor throughput, thus greatly reducing the system performance. As Figure 11 shown in the figure, 4 groups of upstream traffic of L0 enter the link to S0. Assuming that the link bandwidth is 100G each, since the maximum of each group of upstream traffic is 100G, the total of 4 groups is 400G, and the bandwidth of the L0-S0 link is also only 100G, which will result in the actual bandwidth of each group of upstream traffic being only 25G (if the congestion control fairness is not good, there will be a situation where a certain group of traffic is less than 25G), and finally the throughput of the entire system is only 25G (because the traffic topology is a loop, and the throughput of the entire loop is limited by the throughput of the most bottleneck point).

[0141] Even if the network bandwidth is non-convergent, i.e., the total downstream bandwidth of L0 is equal to the total upstream bandwidth, individual link congestion may still occur due to the imperfection of load balancing, ultimately affecting system performance. For example, Figure 11 As shown, assume that the connections of L0-S0 and L0-S1 are both two 100G links. Due to the imperfection of the switch load balancing hash, there is a high probability (90%) that two or more groups of the four groups of upstream traffic will enter the same 100G link. In this case, the throughput of the entire system will drop to 50G or even lower.

[0142] It should be noted that the above discussion is only for a small network as shown in Figure 6 . In an actual large-scale network, the performance randomness is greater, and the unoptimized system performance may deviate more from the optimal performance. For example, in an actual production network, a Rack / Block generally has 32 servers. In the worst case, a Rack will send out 32 groups of cross-Rack traffic. There is a high probability that at least two or more groups of this large amount of traffic will congest a port in the upstream routing hash. In this case, the actual throughput of these congested traffic will drop to 50% or less of the ideal throughput. In the Ring Allreduce operation, as long as there is a bottleneck in the ring, the throughput of the entire operation will be dragged down by this bottleneck, resulting in the completion time of the entire operation increasing to twice or more of the ideal time. For the above reasons, the server node list presented by the manufacturer to the tenant has no topological affinity guarantee. In fact, the node names are often composed of some random letters and numbers, and no topological affinity can be seen from the node names in the list. In addition, the IP (Internet Protocol address) of the node does not have information that can identify the physical location. For the tenant, if the tenant directly uses the node list provided by the cloud provider as the input list of the collective communication library in collective communication, it may result in performance lower than expected. On the other hand, for the cloud provider, due to the randomness of load balancing in the network, even if the network topology is non-convergent in bandwidth design (i.e., the downstream bandwidth and upstream bandwidth of LA are the same), the HPC performance obtained by the tenant will also have randomness (inconsistency) and is difficult to predict.

[0143] Moreover, the above scenarios are all for single-tenant scenarios. In a cloud deployment environment, there are multiple tenants, that is, the servers of multiple tenants are deployed mixedly in the same network. In a multi-tenant scenario, even in the most ideal situation, there will be multiple groups of cross-Rack traffic in the network. If the traffic topologies of multiple tenants are not optimized, there will be more cross-Rack traffic, bringing more performance uncertainties to model training, and then having a non-negligible impact on the SLA (Service-Level Agreement) commitments made by the manufacturer to tenants. For example, as Figure 12 shown, in the scenario of mixed deployment of multi-tenant servers, if there are servers of 2 tenants, among which the servers leased by Tenant 1 include the servers H0, H2, H4, and H6 filled with slashes, and the servers leased by Tenant 2 include the unfilled servers H1, H3, H5, and H7. In the ideal situation, if the traffic topologies of each tenant are optimized, their logical traffic topologies and network traffic are respectively as Figure 13 and Figure 14 shown. The dotted line represents the traffic topology between the servers H0, H2, H4, and H6 of Tenant 1, while the solid line represents the traffic topology between the servers H1, H3, H5, and H7 of Tenant 2. Among them, each access layer switch has 2 groups of uplink traffic, and the probability of congestion of these 2 groups of traffic is relatively small. However, if the traffic topology of each tenant is not optimized, as Figure 15 and Figure 16 shown, each access layer switch has 4 groups of uplink traffic. The probability of congestion of these 4 groups of traffic will be much greater. At the same time, the latency of performing Ring Allreduce operations on the unoptimized topology is also much higher.

[0144] In an actual network, there are 8 Leaf switches and 32 servers in a Rack, and each server has 8 network cards. A Leaf switch has 64 ports for downlink access and also connects 64 ports for uplink. In the extreme case, there will be 64 groups of cross-Rack traffic on the uplink of a Leaf switch. By analyzing and characterizing the relationship between the number of cross-Rack flows and the probability of congestion, as Figure 17As shown, it is a graph of the probability of at least one LA uplink congestion in the access layer switch. The abscissa is the number of uplink flows in the access layer switch, and the ordinate is the probability. In the relationship between the number of uplink flows in the Leaf switch and the probability of congestion, the dashed line is the curve of unoptimized / random, and the solid line is the probability curve of topological optimization and orchestration when a single tenant uses Ring Allreduce. It can be seen that when the number of Leaf uplink flows exceeds 20, even if the switch has 64 uplink links, the probability of at least one link congestion has exceeded 90%. And when the number of uplink flows exceeds 30, there will almost certainly be (at least one link) congestion (the probability is close to 100%).

[0145] For the above practical reasons, in this embodiment, the server node list of the tenant is optimized and orchestrated (such as optimized from Figure 10 as shown to Figure 8 as shown). The optimized traffic topology can minimize the uplink traffic of each rack (such as Figure 8 as shown), thereby minimizing the congestion of the uplink traffic of the access layer switch, improving network smoothness, and maximizing the network throughput of the tenant. In particular, when there is only one tenant in the rack, this embodiment can achieve that there is only one set of uplink traffic in the entire rack, and there will be no congestion in the uplink direction of the access layer switch, and the tenant throughput can reach an ideal level.

[0146] This embodiment can assist the use of the collective communication library in large model training in the form of components. A large model training often has multiple collective operations at the same time, and the servers of multiple collective operations generally do not overlap. For example, in 3-D (Three-Dimensional) parallel training (model parallelism, pipelined parallelism, and data parallelism), all GPUs will participate in an Allreduce operation on one data parallel plane to synchronize the training results of different data parallel planes, and there are dozens of Allreduce operations in total. And each collective communication operation has a master server, which initiates the entire operation. In actual training, as in this solution, the master server of each collective communication operation needs to initiate a topological optimization and orchestration request (the request is completed by the topological optimization proxy component on the server). After the controller receives the request, it will optimize and orchestrate the server list in the request to increase the affinity of the collective communication, reduce the latency and network congestion of the communication operation, improve the throughput of the collective communication, and thus reduce the training time of the large model training.

[0147] The collective communication processing method provided in this embodiment can be implemented in the form of an NCCL Plugin (plugin) by the topology optimization proxy component on the server side when large-scale AI training clusters, such as HCC (High-Performance Computing Cluster) products, are implemented. The controller is implemented by a dedicated server and can specifically connect to the corresponding network topology database. As Figure 18 shown, for different collective operations, each collective operation is implemented by multiple servers, including a master server and several slave servers. A collective communication library, specifically NCCL, can be set in the master server. The master server can send the original node list to the controller, and the controller returns the optimal node list after arranging and updating the original node list, so that the master server can initiate the corresponding collective operation based on the received optimal node list.

[0148] Specifically, in actual business, a tenant will obtain a set of original server node lists after purchasing servers. During model training, the tenant uses a training framework (such as TensorFlow or PyTorch, etc.) to allocate server resources. When performing a certain collective operation, the training framework will pass part of the node list to the collective communication library (such as NCCL or MPI, etc.). Generally, in model training, not all servers are used for the same collective operation. Each collective operation is only a part of the training, and there are many collective operations in the whole training, and each collective operation only uses some servers. When the training framework calls the collective communication library to complete the collective operation, it will pass the server list (operation node list) used for this operation to the collective communication library. As Figure 19 shown, in step 0, the training framework of the master server obtains all the original node lists and passes the operation node list as the original node list to the collective communication library.

[0149] Furthermore, the composition of this embodiment mainly includes: a topology optimization proxy on the server side and a centralized controller providing topology sorting capabilities. The controller can be a physically independent server or a logical module on an existing server. The proxy and the controller communicate in a "request-response" manner through the network. In this embodiment, the collective communication library can be slightly modified to add a request topology optimization interface. In some communication libraries, instead of modifying the communication library itself, its plugin component can be modified. For example, NCCL has a built-in plugin function interface, and only this interface needs to be added to the NCCL plugin.

[0150] As Figure 19In step ①, when the collective communication library of the master server receives a collective communication operation command, it will first request topology orchestration optimization from the topology optimization agent of this solution; in step ②, after receiving the request from the collective communication library, the topology optimization agent will request a topology sorting request from the controller through the network; in step ③, after receiving the request, the topology sorting engine on the controller will first query the topology database; in step ④, the controller sorts according to the topology information of each server node, and then sends the sorted server list back to the topology optimization agent; in step ⑤, after receiving the topology orchestration optimization response, the topology optimization agent will return the result to the collective communication library of the master server, and the collective communication library will initiate collective communication using the orchestrated server list.

[0151] Furthermore, for topology clustering, that is, affinity-aware processing. The topology sorting engine first queries the topology information of each server in the operation server list from the topology database. The form of the topology information is similar to {HostId:ClusterId, PodId, RackId}, and the topology information can be queried in the topology database using the IP or fixed asset number of the server, etc. As Figure 20 shown, each server node can be divided according to Cluster, Pod, and Rack, and the topology information of server node H16 can be expressed as {H16: 1,2,4}.

[0152] After querying the topology information of all servers, the topology sorting engine will perform hierarchical clustering on the servers. The specific steps are as follows:

[0153] a) Cluster all servers according to ClusterId, that is, gather the servers with the same ClusterId together (arranged adjacent to each other). The actual operation is to traverse all servers, check the ClusterId of the servers. If there is a class with the ClusterId in the clustering, add the server to this class; otherwise, create a new class with the ClusterId and add the server to this class.

[0154] b) In each clustering with the same ClusterId, cluster according to PodId, that is, gather the servers with the same PodId together (arranged adjacent to each other). The actual operation is to traverse all servers in each Cluster class, check the PodId of the servers. If there is a class with the PodId in the clustering, add the server to this class; otherwise, create a new class with the PodId and add the server to this class.

[0155] c) In each cluster with the same ClusterId and PodId, cluster according to RackId, that is, gather the servers with the same PodId together (arranged adjacent to each other). The actual operation is to traverse all the servers in each Pod class, check the RackId of the servers. If there is a class with the RackId in the cluster, add the server to this class; otherwise, create a new class with the RackId and add the server to this class. So far, all the servers participating in this collective communication operation have been clustered, that is, affinity awareness has been achieved.

[0156] As Figure 20 shown, the list of servers participating in a certain collective communication operation in each server may be: [H11, H25, H14, H21, H2, H6, H13, H0, H22, H30, H16, H5, H29, H23, H7, H18]. Then, as Figure 21 shown, for the Figure 20 clustering result of the list of servers participating in a certain collective communication operation, specifically, after clustering by ClusterId, 2 Cluster classes are obtained, namely {0: [H11, H14, H2, H6, H13, H0, H5, H7]} and {1: [H25, H21, H22, H30, H16, H29, H23, H18]}. After clustering by PodId, 4 Pod classes can be obtained, namely {0: [H2, H6, H0, H5, H7]}, {1, [H11, H14, H13]}, {2: [H21, H22, H16, H23, H18]} and {3: [H25, H30, H29]}. After clustering by RackId, 8 Rack classes can be obtained {0: [H2, H0]}, {1: [H6, H5, H7]}, {2: [H11]}, {3: [H14, H13]}, {4: [H16, H18]}, {5: [H21, H22, H23]}, {6: [H25]} and {7: [H30, H29]}.

[0157] Furthermore, for node orchestration optimization. After obtaining the clustering result, it is necessary to perform topology optimization orchestration based on the result. The specific method is to perform a depth-first search / traversal on the clustering result, that is, in the clustering tree, search for the Pods under each Cluster one by one, and then search for the Racks under each Pod one by one, and put the searched nodes into the optimization list L in turn.

[0158] For example, for the Figure 21 shown clustering tree:

[0159] a) First search Cluster 0. Since it has not reached the leaf node, further search Pod 0. Since it has not reached the leaf node, further search Rack 0. Rack 0 is already a leaf node. Put its server (H2, H0) into the optimization list L. At this time, L=[H2, H0].

[0160] b) Then search for other Racks under Pod 0, find Rack 1, and put its servers into the optimization list L. At this time, L=[H2,H0,H6,H5,H7].

[0161] c) At this point, the leaves of Pod 0 have been searched, and the other Pods in Cluster 0 are searched, that is, Pod 1. Rack 2 below it is searched, and its servers are placed in L, where L = [H2, H0, H6, H5, H7, H11].

[0162] d) Then search for Rack 3 under Pod 1 and put its servers into L, where L = [H2, H0, H6, H5, H7, H11, H14, H13].

[0163] e) At this point, Cluster 0 has been searched, and Cluster 1 is searched next. This process continues until all racks are searched, and the final optimization list is L=[H2,H0,H6,H5,H7,H11,H14,H13,H16,H18,H21,H22,H23,H25,H30,H29].

[0164] Note that in the above process, the search order is not unique. For example, you can search Cluster 1 first and then Cluster 0. For another example, in Pod 0, you can search Rack 1 first and then Rack 0. In addition, the arrangement of each server in a Rack is also not restricted in order. For example, in Rack 0, it is the same whether H0 or H2 is arranged first, and the final logical effect is equivalent.

[0165] When the arranged list is used for Ring Allreduce operation, the traffic topology formed is as follows Figure 22As shown, where different heights of the connections represent the switch levels crossed by the traffic of this group. For example, H13->H16 and H29->H2 both cross the backbone layer switches, and H7->H11 and H23->H25 both cross the core layer switches. It can be seen that the traffic after optimized orchestration has minimized the number of flows crossing the aggregation layer, core layer, and backbone layer switches. Thus, the probability of congestion of these flows with themselves and other flows can be minimized, thereby improving the training throughput. It should be noted that the optimized orchestration topology in this embodiment can also be used for other implementations of Allreduce, including implementations such as Tree-Allreduce and Halving-Doubling.

[0166] In addition, the collective communication processing method provided in this embodiment can also vary the search starting point. In this embodiment, the search starts from the 0th Rack of the 0th Pod in the 0th Cluster. As an example of variation, it can also start from the 4th Rack of the 2nd Pod in the 1st Cluster. The ultimately optimized orchestration topology ring is actually the same, just rotated, that is, the starting point of the server nodes is different. The order of nodes within the same Rack can also be varied. In this embodiment, for multiple nodes in the same Rack, they are arranged in the order in the original operation node list, but actually they can be rearranged arbitrarily.

[0167] The collective communication processing method provided in this embodiment optimizes the traffic scheduling of collective operations in large language model training (AI HPC), including the traffic scheduling of the most widely used Allreduce operation. This embodiment uses the technology of centralized topological sorting and a request-response-based architecture to optimize the orchestration of the server list provided by cloud HPC service providers to tenants, improving the affinity of collective operations, thereby optimizing the traffic scheduling and throughput performance of collective communication. This embodiment can eliminate congestion of HPC traffic in the network in a single-tenant scenario; and in a multi-tenant scenario, it can also minimize network congestion, thereby maximizing the throughput and efficiency of the HPC cluster, while significantly improving the smoothness and stability of the network. Compared with the unoptimized scenario, this embodiment can provide a throughput increase of up to 70%.

[0168] The collective communication processing method provided in this embodiment can minimize the latency of collective communication operations. Especially in the Ring Allreduce operation, the optimized topology orchestration can reduce the communication latency by more than 50% compared to the unoptimized orchestration. The collective communication processing method provided in this embodiment can also minimize network bandwidth competition (network congestion), improve network smoothness, eliminate the bottleneck of the traffic loop in collective communication, and maximize the throughput of collective operations. Specifically, if there is only one tenant in the network, this embodiment can completely eliminate congestion and make the network throughput reach the ideal value. AsFigure 23 As shown in the figure, in the 100Gb network single-tenant scenario, 10 servers distributed in 2 Racks are used for Ring Allreduce operation. After each of the two schemes runs 200 times on a 100Gb network, the optimized and orchestrated topology can fully utilize the system throughput, and the net throughput (Busbw) reaches about 12GB; while the throughput achieved by the unoptimized and orchestrated traffic topology is often only about 7-8GB. If there are multiple tenants, the throughput lost due to network congestion in the unoptimized and orchestrated traffic may be even more.

[0169] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0170] Based on the same inventive concept, the embodiments of the present application also provide a collective communication processing device for implementing the collective communication processing method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the collective communication processing device provided below can refer to the limitations on the collective communication processing method in the above text, and will not be repeated here.

[0171] In an exemplary embodiment, as Figure 24 shown, a collective communication processing device 2400 is provided, including: a node orchestration request acquisition module 2402, a topology information acquisition module 2404, a node orchestration update module 2406, and a collective operation processing module 2408, where:

[0172] The node orchestration request acquisition module 2402 is configured to acquire a node orchestration request for a target collective operation in collective communication; the node orchestration request carries an original node list, and the original node list includes each server node for the target collective operation;

[0173] The topology information acquisition module 2404 is configured to acquire the topology information of each server node; the topology information is used to describe the switching nodes connected by the server nodes at each level in the node communication network; each server node communicates based on the switching nodes connected at each level;

[0174] A node arrangement update module 2406, configured to, based on each piece of topology information, arrange and update the arrangement order of each server node level by level according to the levels in the node communication network, so as to obtain an arranged node list;

[0175] A set operation processing module 2408, configured to perform set communication for a target set operation through each server node according to the arrangement order of each server node in the arranged node list.

[0176] In one embodiment, the node arrangement update module 2406 is further configured to, based on each piece of topology information, hierarchically cluster each server node level by level according to the levels in the node communication network, so as to obtain a topology clustering result of each server node; based on each topology clustering result, perform node search level by level according to the levels in the node communication network to obtain an arranged node list; the arrangement order of each server node in the arranged node list conforms to the hierarchical order of the node search.

[0177] In one embodiment, the node arrangement update module 2406 is further configured to, based on each piece of topology information, hierarchically cluster each server node level by level according to the hierarchical dimension from high to low in the node communication network, so as to obtain a topology clustering result of each server node; in the node communication network, a high-level switching node is used to support cross-node communication between server nodes connected by a low-level switching node.

[0178] In one embodiment, the levels in the node communication network include a core layer, an aggregation layer, and an access layer from high to low;

[0179] In one embodiment, the node arrangement update module 2406 is further configured to cluster each server node based on each piece of topology information according to the hierarchical dimension of the core layer to obtain a core layer clustering result of each server node; cluster each server node based on each core layer clustering result according to the hierarchical dimension of the aggregation layer to obtain an aggregation layer clustering result of each server node; cluster each server node based on each aggregation layer clustering result according to the hierarchical dimension of the access layer to obtain a topology clustering result of each server node.

[0180] In one embodiment, the node arrangement update module 2406 is further configured to, based on each topology clustering result, perform node search for each server node level by level according to the hierarchical dimension from high to low in the node communication network to obtain an arranged node list; in the node communication network, a high-level switching node is used to support cross-node communication between server nodes connected by a low-level switching node.

[0181] In one embodiment, the levels in the node communication network include a core layer, an aggregation layer, and an access layer from high to low; the node orchestration update module 2406 is further configured to, based on each topology clustering result, perform node search for each server node step by step according to the level dimension from high to low in the node communication network to obtain an orchestration node list, including: performing search according to the level dimension of the core layer based on each topology clustering result to obtain a core layer search result; performing search according to the level dimension of the aggregation layer based on the core layer search result to obtain an aggregation layer search result; performing search according to the level dimension of the access layer based on the aggregation layer search result to obtain an access layer search result; and obtaining the orchestration node list according to the core layer search result, the aggregation layer search result, and the access layer search result.

[0182] In one embodiment, the node orchestration request acquisition module 2402 is further configured to determine a collective communication library in collective communication; and obtain a node orchestration request for a target collective operation in collective communication based on a communication interface with the collective communication library.

[0183] In one embodiment, the topology information acquisition module 2404 is further configured to determine a topology database associated with the node communication network; and query the topology information of each server node from the topology database according to the original node list.

[0184] In one embodiment, the node orchestration request acquisition module 2402 is further configured to obtain a node orchestration request from a target server node; the target server node is determined from each server node for a target collective operation in collective communication; the collective operation processing module 2408 is further configured to send the orchestration node list to the target server node; and the orchestration node list is used to instruct the target server node to perform collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestration node list.

[0185] In one embodiment, it further includes a trigger determination module, configured to execute the step of obtaining a node orchestration request for a target collective operation in collective communication when each server node for the target collective operation meets an orchestration update trigger condition.

[0186] In an exemplary embodiment, as Figure 25 shown, a collective communication processing apparatus 2500 is provided, including: an original node list acquisition module 2502, a node orchestration request module 2504, and a collective operation processing module 2506, where:

[0187] The original node list acquisition module 2502 is configured to determine an original node list when a target collective operation in collective communication is triggered; the original node list includes each server node for the target collective operation;

[0188] A node orchestration request module 2504 is configured to send a node orchestration request generated according to an original node list to a server; the node orchestration request is used to instruct the server to return an orchestrated node list; the orchestrated node list is obtained by orchestrating and updating the arrangement order of each server node in the original node list level by level based on the topology information of each server node according to the levels in the node communication network.

[0189] A set operation processing module 2506 is configured to receive the orchestrated node list and perform set communication for a target set operation through each server node according to the arrangement order of each server node in the orchestrated node list.

[0190] Each module in the above set communication processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in a processor in a computer device in hardware form or be independent of the processor, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0191] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 26 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store set communication processing data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a set communication processing method. Those skilled in the art can understand that Figure 26 the structure shown in

[0192] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout. In an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0193] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0194] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0195] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0196] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0197] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification. The above embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A collective communication processing method, characterized in that: The method comprises: Obtaining a node orchestration request for a target set operation in the set communication; the node orchestration request carries an original node list, and the original node list includes each server node for the target set operation; Acquire topology information of each of the server nodes; the topology information is used to describe the switch nodes connected to the server nodes at each level of the node communication network; each of the server nodes communicates with each other based on the switch nodes connected at each level; Based on each of the topology information, the arrangement order of each of the server nodes is updated step by step according to the levels in the node communication network to obtain an arrangement node list; According to the arrangement order of each of the server nodes in the orchestration node list, collective communication is performed for the target collective operation through each of the server nodes.

2. The method according to claim 1, characterized in that The step of arranging and updating the arrangement order of each of the server nodes step by step based on each of the topology information and according to the levels in the node communication network to obtain an arrangement node list includes: Based on each of the topology information, hierarchically clustering each of the server nodes according to the levels in the node communication network, to obtain a topological clustering result of each of the server nodes; Based on each of the topological clustering results, node search is performed level by level according to the hierarchy in the node communication network to obtain an orchestration node list; the arrangement order of each of the server nodes in the orchestration node list conforms to the hierarchical order of the node search.

3. The method according to claim 2, characterized in that Based on each of the topology information, hierarchically clustering each of the server nodes step by step according to the levels in the node communication network to obtain a topological clustering result of each of the server nodes includes: Based on each of the topology information, hierarchically clustering each of the server nodes step by step according to the hierarchical dimensions from high to low in the node communication network, to obtain a topological clustering result of each of the server nodes; In the node communication network, high-level switching nodes are used to support cross-node communication between server nodes connected to low-level switching nodes.

4. The method according to claim 3, characterized in that The layers in the node communication network include, from high to low, a core layer, a convergence layer, and an access layer; Based on each of the topology information, hierarchically clustering each of the server nodes step by step according to the hierarchical dimensions from high to low in the node communication network to obtain a topological clustering result of each of the server nodes, including: Based on each of the topology information, clustering each of the server nodes according to the hierarchical dimension of the core layer to obtain a core layer clustering result of each of the server nodes; Based on the clustering results of each core layer, clustering each server node according to the hierarchical dimension of the aggregation layer to obtain the aggregation layer clustering results of each server node; Based on the clustering results of each convergence layer, each server node is clustered according to the hierarchical dimension of the access layer to obtain a topological clustering result of each server node.

5. The method according to claim 2, characterized in that: The step of searching nodes level by level according to the levels in the node communication network based on the topological clustering results to obtain an orchestration node list includes: Based on the topological clustering results, according to the hierarchical dimensions from high to low in the node communication network, node searches are performed for each of the server nodes level by level to obtain an orchestration node list; In the node communication network, high-level switching nodes are used to support cross-node communication between server nodes connected to low-level switching nodes.

6. The method according to claim 5, characterized in that The layers in the node communication network include, from high to low, a core layer, a convergence layer, and an access layer; Based on each of the topological clustering results, according to the hierarchical dimensions from high to low in the node communication network, node searches are performed level by level for each of the server nodes to obtain an orchestration node list, including: Based on each of the topological clustering results, searching is performed according to the hierarchical dimension of the core layer to obtain core layer search results; Based on the core layer search results, a search is performed according to the hierarchical dimension of the aggregation layer to obtain the aggregation layer search results; Searching according to the hierarchical dimension of the access layer based on the convergence layer search results to obtain access layer search results; An orchestration node list is obtained according to the core layer search result, the convergence layer search result, and the access layer search result.

7. The method according to claim 1, characterized in that The obtaining of a node orchestration request for a target set operation in the set communication includes: Determining a collective communication library in the collective communication; Based on the communication interface with the collective communication library, a node orchestration request for a target collective operation in the collective communication is obtained.

8. The method according to claim 1, characterized in that The obtaining of topology information of each of the server nodes includes: A topological database for determining node communication network associations; According to the original node list, the topology information of each server node is queried from the topology database.

9. The method according to claim 1, characterized in that: The obtaining of a node orchestration request for a target set operation in the set communication includes: Obtaining a node scheduling request from a target server node; the target server node is determined from each server node for a target set operation in a set communication; The performing collective communication for the target collective operation through each of the server nodes according to the arrangement order of each of the server nodes in the orchestration node list includes: The orchestration node list is sent to the target server node; the orchestration node list is used to instruct the target server node to perform collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestration node list.

10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: When each server node operating on the target set meets the orchestration update triggering condition, the step of obtaining the node orchestration request for the target set operation in the collective communication is performed.

11. A collective communication processing method, characterized in that: The method comprises: When a target set operation in the set communication is triggered, an original node list is determined; the original node list includes each server node for the target set operation; Sending a node orchestration request generated according to the original node list to the server; the node orchestration request is used to instruct the server to return an orchestration node list; the orchestration node list is obtained by arranging and updating the arrangement order of each server node in the original node list step by step according to the hierarchy in the node communication network based on the topology information of each server node, and the topology information is used to describe the switching nodes connected to the server node in each hierarchy of the node communication network; An orchestration node list is received, and according to the arrangement order of each of the server nodes in the orchestration node list, collective communication is performed through each of the server nodes for the target collective operation.

12. A collective communication processing device, characterized in that: The device comprises: A node orchestration request acquisition module, used to acquire a node orchestration request for a target set operation in a set communication; the node orchestration request carries an original node list, and the original node list includes each server node for the target set operation; A topology information acquisition module, used to acquire topology information of each of the server nodes; the topology information is used to describe the switch nodes connected to the server nodes in each level of the node communication network; each of the server nodes communicates with each other based on the switch nodes connected in each level; A node arrangement update module, configured to arrange and update the arrangement order of each server node step by step according to the hierarchy in the node communication network based on each of the topology information, so as to obtain an arrangement node list; The collective operation processing module is used to perform collective communication for the target collective operation through each of the server nodes according to the arrangement order of each of the server nodes in the orchestration node list.

13. A collective communication processing device, characterized in that: The device comprises: The original node list acquisition module is used to determine the original node list when the target set operation in the set communication is triggered; the original node list includes each server node for the target set operation; A node orchestration request module, configured to send a node orchestration request generated according to the original node list to a server; the node orchestration request is used to instruct the server to return an orchestration node list; the orchestration node list is obtained by orchestrating and updating the arrangement order of each server node in the original node list step by step based on the topology information of each server node and according to the level in the node communication network, and the topology information is used to describe the switching nodes connected to the server node in each level of the node communication network; The collective operation processing module is used to receive the orchestration node list, and perform collective communication for the target collective operation through each server node according to the arrangement order of each server node in the orchestration node list.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • System and method for supporting partition-aware routing in a multi-tenant cluster environment

    CN107113233A

  • Route acquisition method and device of software-defined network, and storage medium

    CN110572323A