A method, apparatus, cluster, program product, and medium for aggregated communication
Patent Information
- Application Number
- CN202510293045.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]然而,受限于显存资源,通信缓存区(buffer)通常不会配置的很大,在跨数据中心进行AllReduce集合通信操作的过程中,需要分多轮通信才能完成,由此会带来巨大的链路传播时延
[0026]第五方面,本申请实施例提供了一种计算机可读存储介质,计算机可读存储介质中存储有至少一条计算机程序,计算机程序由处理器加载并执行以实现如上述第一方面的集合通信方法。
Smart Images

Figure CN122802588A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of aggregated communication technology, and in particular to an aggregated communication method, apparatus, cluster, program product and medium. Background Technology
[0002] Due to constraints in physical resources such as electricity, the computing power of a single data center (DC) cannot grow indefinitely. This means that some large models cannot be deployed within the computing resources of a single data center. To effectively utilize the fragmented computing resources that may exist across multiple data centers, the industry has begun to explore joint training of large models across multiple data centers. Cross-data center model training typically employs pipeline parallelism (PP), data parallelism (DP), or a hybrid of both, involving communication such as cross-data center reduction operations (AllReduce) communication.
[0003] Currently, there is no collective communication library specifically for cross-data center scenarios. The Huawei Collective Communication Library (HCCL) or the NVIDIA Collective Communication Library (NCCL) are typically used.
[0004] However, due to limitations in video memory resources, the communication buffer is usually not configured to be very large. During the AllReduce collection communication operation across data centers, multiple rounds of communication are required to complete the operation, which will result in huge link propagation latency. Summary of the Invention
[0005] This application provides a collection communication method, apparatus, cluster, program product, and medium that can improve the utilization rate of the communication buffer and further reduce the latency of collection communication operations.
[0006] Firstly, a collective communication method is provided for application in a first computing cluster, the method comprising:
[0007] Obtain the target data indicated by the target communication task; store the target data in the communication buffer between the first computing cluster and the second computing cluster; and, if the amount of data in the communication buffer meets the preset conditions, perform a collective communication operation with the second computing cluster on the target data in the communication buffer.
[0008] As shown above, in order to reduce communication latency across computing clusters during inter-cluster communication, it is necessary to minimize inter-cluster aggregate communication operations. Therefore, improving the utilization rate of the communication buffer is extremely important. In the method described above, when the amount of data in the communication buffer meets the preset conditions, the first computing cluster and the second computing cluster perform aggregate communication operations on the target data in the communication buffer, effectively improving the utilization rate of the communication buffer, reducing the number of inter-cluster aggregate communications, and thus reducing latency during aggregate communication.
[0009] In one possible implementation, the target data is the data to be transmitted as indicated by the target communication task, or the result data of at least one round of cluster-wide collective communication operations performed between computing nodes in the first computing cluster for the data to be transmitted.
[0010] As shown above, when the target communication task includes an AllGather operation, the target data is the data to be transmitted as indicated by the target communication task. This is because the amount of data increases round by round during the AllGather operation. Adopting a layered communication approach—first inter-cluster communication, then intra-cluster communication—reduces the number of cross-cluster aggregate communication operations, thereby reducing latency. When the target communication task includes ReduceScatter, the target data is the result data of at least one round of intra-cluster aggregate communication operations performed between computing nodes in the first computing cluster for the data to be transmitted. This is because the amount of data decreases round by round during the ReduceScatter operation. Adopting a layered communication approach—first intra-cluster communication, then inter-cluster communication—reduces the number of cross-cluster aggregate communication operations, thereby reducing latency.
[0011] In one possible implementation, the preset conditions include the data capacity of the data to be transmitted in the communication buffer being a first preset threshold; or the data volume of the communication buffer not being increased within a preset time period; the first preset threshold is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain.
[0012] As shown above, when the target communication task includes an AllGather operation, the preset condition includes a first preset threshold for the data capacity of the data to be transmitted in the communication buffer. This first preset threshold is the ratio of the communication buffer capacity to the number of computing clusters in the communication domain. This ensures that the communication buffer is sufficient to store the result data of the AllGather operation, thereby reducing the number of cross-cluster communications. When the target communication task includes ReduceScatter, the preset condition includes that the data volume of the communication buffer capacity does not increase within a preset time period. Specifically, this could mean that the utilization rate of the communication buffer capacity reaches its maximum, or that there is no result data of at least one round of intra-cluster aggregate communication operation between computing nodes in the first computing cluster for the data to be transmitted within the preset time period. This improves the utilization rate of the communication buffer and reduces the number of cross-cluster communications.
[0013] In one possible implementation, the first computing cluster and the second computing cluster belong to a first communication domain, and the capacity of the communication buffer is greater than the capacity of the communication buffer of the second communication domain; the computing nodes in the second communication domain belong to at most one computing cluster.
[0014] As can be seen from the above, different configurations of the communication buffer size for different communication scenarios can make reasonable use of video memory and improve the utilization rate of the communication buffer.
[0015] In one possible implementation, if the target communication task includes ReduceScatter, a collective communication operation between computing nodes in the first computing cluster is performed to obtain the target data based on the data to be transmitted as indicated by the target communication task.
[0016] As shown above, when the target communication task includes ReduceScatter, the target data is the result of at least one round of intra-cluster aggregate communication operations between computing nodes in the first computing cluster for the data to be transmitted. This is because the amount of data decreases round by round during the ReduceScatter operation. By adopting a hierarchical communication approach—first within the cluster, then between clusters—the number of cross-cluster aggregate communications can be reduced, thereby reducing latency.
[0017] In one possible implementation, the data to be transmitted indicated by the target communication task is stored in the communication buffer according to the capacity of the communication buffer, so that the capacity of the communication buffer is utilized to the maximum preset utilization rate.
[0018] As can be seen from the above, when the target communication task includes ReduceScatter, when the first computing cluster performs the ReduceScatter operation within the cluster, it stores the data to be transmitted indicated by the target communication task into the communication buffer according to the capacity of the communication buffer, so that the utilization rate of the communication buffer reaches the maximum preset utilization rate, thereby improving the utilization rate of the communication buffer and reducing latency.
[0019] In one possible implementation, when the target communication task includes AllGather, the target data is the data to be transmitted as indicated by the target communication task; the data volume of the target data is determined; the data volume of the target data is a first preset threshold; the first preset threshold is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain.
[0020] As shown above, when the target communication task includes an AllGather operation, the target data is the data to be transmitted as indicated by the target communication task. This is because the amount of data increases round by round during the AllGather operation, requiring a hierarchical communication approach: first between clusters, then within clusters. The target data volume is the ratio of the communication buffer capacity to the number of computing clusters in the communication domain, which reduces the number of cross-cluster aggregate communications, thereby reducing latency.
[0021] In one possible implementation, when the target communication task includes AllGather, the result data of the collective communication operation with the second computing cluster on the target data in the communication buffer is stored in the communication buffer, such that the data volume of the communication buffer is a second preset threshold; the second preset threshold is the ratio of the capacity of the communication buffer to the number of computing nodes in the first computing cluster; based on the result data of the collective communication operation in the communication buffer, the computing nodes in the first computing cluster perform collective communication operations.
[0022] As shown above, when the target communication task includes an AllGather operation, a hierarchical communication approach is adopted, first between clusters and then within clusters. After the first and second computing clusters perform a set communication operation on the communication buffer, the result data of the set communication operation between the clusters and the second computing cluster on the target data in the communication buffer is stored in the communication buffer. This ensures that the data volume in the communication buffer is at a second preset threshold. The AllGather operation is then performed between computing nodes in the first computing cluster to guarantee that the communication buffer has sufficient space to store the result data of the AllGather operation, thereby improving the utilization rate of the communication buffer.
[0023] Secondly, a collective communication device is provided. Embodiments of this application can divide the collective communication device into functional modules according to the method provided in the first aspect. For example, each function can be divided into its own functional modules, or two or more functions can be inherited into a single processing module. For instance, embodiments of this application can divide the collective communication device into an acquisition module, a caching module, and a communication module according to their functions. Descriptions of the possible technical solutions and beneficial effects of the various functional modules described above can be found in the technical solutions provided in the first aspect or its corresponding possible implementations, and will not be repeated here.
[0024] Thirdly, embodiments of this application provide a computing cluster, which includes at least one computing device, each computing device including a processor and a memory for storing processor-executable instructions; the processor is configured to execute instructions, causing the computing device to perform the set communication method provided in the first aspect above.
[0025] Fourthly, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computing node reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing node to perform the collective communication method provided in the various optional implementations of the first aspect described above.
[0026] Fifthly, embodiments of this application provide a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to implement the collection communication method as described in the first aspect above.
[0027] For a detailed description of the second to fifth aspects and their various implementations in the embodiments of this application, please refer to the detailed description in the first aspect and its various implementations; and for a detailed description of the beneficial effects of the second to fifth aspects and their various implementations, please refer to the beneficial effect analysis in the various implementations of the first aspect, which will not be repeated here.
[0028] These or other aspects of the embodiments of this application will become more apparent in the following description. Attached Figure Description
[0029] Figure 1 This diagram illustrates the use of buffers in data flow during communication in NCCL, as provided by related technologies.
[0030] Figure 2 This diagram illustrates a data stream buffer usage scenario in HCCL communication provided by related technologies.
[0031] Figure 3 A schematic diagram of the hardware architecture of a collective communication system 100 provided in an embodiment of this application is shown;
[0032] Figure 4 This paper illustrates a software architecture diagram of a collective communication system 100 provided in an embodiment of this application;
[0033] Figure 5 A flowchart illustrating a collection communication method provided in an embodiment of this application is shown;
[0034] Figure 6 This illustration shows a communication process diagram of a ReduceScatter operation provided in an embodiment of this application;
[0035] Figure 7 This illustration shows a communication process diagram of an AllGather operation provided in an embodiment of this application;
[0036] Figure 8 A schematic diagram of the structure of a collection communication device provided in an embodiment of this application is shown. Detailed Implementation
[0037] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish the same or similar items with basically the same function and effect.
[0038] Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or order of execution, and that "first," "second," etc., are not necessarily different. Furthermore, in some embodiments of this application, words such as "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.
[0039] Furthermore, the device architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of device architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0040] Terminology introduction:
[0041] The message passing interface (MPI) is a standardized and widely used parallel programming model for developing programs that can run in parallel on multiple computing nodes (e.g., clusters, multi-core systems).
[0042] A communication domain is a mechanism for grouping processes that participate in MPI communication. When an MPI program runs, it often starts numerous processes, and communication domains divide these processes into different groups. Processes within the same communication domain can communicate and collaborate with each other; different communication domains are isolated from each other and do not interfere with each other.
[0043] Collective communication primitives are instructions used in parallel and distributed computing environments to coordinate data transfer, synchronization, and other collective operations among multiple processes or threads. Common collective communication primitives include broadcast, reduce, scatter, and gather. Broadcast means sending the same data from one process to all other processes in the communication domain. Reduce means that all processes in the communication domain participate in a computation (e.g., summation, product, maximum / minimum), merging their respective data into a single final result, which is usually stored in the root process. Scatter means sending different data from one process to all different processes in the communication domain. For example, splitting the root process's data block according to predetermined rules and distributing it to different processes. Gather means collecting data from various processes in the communication domain into a specified process.
[0044] First, the application scenarios of the embodiments of this application will be introduced by way of example.
[0045] A data center is a facility that centrally houses computer systems and related components, typically including a large number of servers, storage devices, and network equipment. Its primary function is to store data and provide computing power services. Limited by basic resources such as electricity, the computing power of a single data center cannot grow indefinitely, making it impossible to deploy large models within the resources of a single data center. To achieve the deployment of large models and effectively utilize the fragmented computing resources of data centers, attempts have begun to be made to jointly train large models using the computing resources of multiple data centers. Cross-data center model training generally involves AllReduce communication in aggregated communication, typically using NCCL or HCCL. AllReduce is a commonly used communication operation in distributed computing environments. In a group of computing nodes, each node has its own process and corresponding data. The purpose of the AllReduce operation is to perform communication and aggregation operations between these computing nodes, ensuring that all nodes ultimately obtain the same result. Specifically, the AllReduce operation first performs a certain operation (such as summation or averaging) on the data on each computing node, and then broadcasts the aggregated value to all computing nodes.
[0046] Taking a ring topology as an example for a data center that implements All Reduce communication, Figure 1 This diagram illustrates the use of buffers in data flow during communication in NCCL, as provided by related technologies. Figure 1 As shown, the raw data is distributed to N node Rings through different data channels. Here, N is a positive integer greater than 3. A data channel is a path or mechanism used for data transmission in fields such as data processing and communication. Specifically, when processing the data generated by AllReduce, each data channel divides the AllReduce data into multiple data blocks and performs send / receive operations in multiple rounds. During a single round of sending data block 1, data block 1 is typically copied multiple times to buffer 1 and sent / received through a send / receive queue.
[0047] Figure 2 This diagram illustrates a data stream buffer usage scenario in HCCL communication, as provided by related technologies. For example... Figure 2As shown, the original data is divided into multiple executions based on the size of the buffer in the HCCL communication library: execution 1, execution 2, and execution 3. The number of executions depends on the size of the original data and the size of the buffer. Before the collection communication task begins, the data from execution 1 is copied to buffer 1. This data is further divided into multiple data slices, and each task step processes one data slice, performing send / receive operations through the send / receive queue.
[0048] As mentioned above, in the NCCL implementation, the amount of data that can be sent in a single batch depends on the size of the buffer and the number of data channels. However, due to limitations in GPU memory resources, the buffer is generally not configured to be very large, with the default configuration showing a single data transmission size in the tens of MB range. During the training of large models across data centers, the AllReduce data for gradient synchronization is in the GB range, causing the AllReduce task to require multiple rounds to complete. Similarly, in the HCCL implementation, the buffer size is also limited, and gradient synchronization still requires multiple rounds to complete. This results in significant link propagation latency. When performing aggregated communication across data centers, the round-trip time (RRT) is typically required to be in the millisecond range, while the latency caused by the aforementioned multiple rounds of cross-data center transmission is usually in the second range. Therefore, NCCL or HCCL implementations are not suitable for scenarios involving training large models across data centers.
[0049] In view of this, embodiments of this application provide a collective communication method applied to a first computing cluster. The method includes: acquiring target data indicated by a target communication task; storing the target data in a communication buffer between the first and second computing clusters; and, when the amount of data in the communication buffer meets a preset condition, performing a collective communication operation with the second computing cluster on the target data in the communication buffer. As can be seen, during the cross-cluster collective communication between the first and second computing clusters, the target data indicated by the target communication task is stored in the communication buffer; when the amount of data in the communication buffer meets a preset condition, the first and second computing clusters perform a cross-cluster communication collective operation on the target data in the communication buffer. The second computing cluster is a computing cluster other than the first computing cluster. This method can improve the utilization rate of the communication buffer, reduce the number of cross-cluster collective communications, and thus effectively reduce the link propagation latency during the collective communication process.
[0050] Secondly, the system architecture of the embodiments of this application will be described by way of example.
[0051] Figure 3A schematic diagram of the hardware architecture of a collective communication system 100 provided in an embodiment of this application is shown. Figure 3 As shown, the aggregated communication system 100 includes a first computing cluster 110, a second computing cluster 120, a switch 130, and a communication buffer 140.
[0052] Specifically, the first computing cluster 110 and the second computing cluster 120 can be a first data center and a second data center. A data center is a facility that centrally stores computer systems and related components, typically including a large number of servers, storage devices, network equipment, etc. Its main purpose is to store data and provide computing power services.
[0053] The first computing cluster 110 and the second computing cluster 120 each include multiple computing nodes. The computing nodes are used to perform aggregated communication operations, processing, receiving, or sending target data indicated by the target communication task. The computing nodes can be computing devices, specifically standard general-purpose servers, such as blade servers, high-density servers, rack servers, or high-performance servers. Optionally, the computing nodes can be graphics processing units (GPUs).
[0054] The first computing cluster 110 is used to acquire the target data indicated by the target communication task and store the target data in the communication buffer 140 between the first computing cluster and the second computing cluster. When the amount of data in the communication buffer 140 meets the preset conditions, the first computing cluster 110 and the second computing cluster 120 perform a set communication operation on the target data in the communication buffer 140.
[0055] Communication buffer 140 is a memory area used to store target data.
[0056] Switch 130 is a network device used for forwarding electrical signals. In a computing cluster, switch 130 is used to connect the devices corresponding to various computing nodes together, building an interconnected network environment that enables the computing nodes to communicate with each other and work collaboratively.
[0057] Taking the scenario of training artificial intelligence models across computing clusters as an example, Figure 4 A software architecture diagram of a collective communication system 100 provided in an embodiment of this application is shown. Figure 4 As shown, the software architecture of the aggregated communication system 100 includes an application layer 210, an architecture layer 220, an aggregated communication layer 230, and a driver layer 240.
[0058] The application layer 210 is used to run the input data for the collection communication, which may specifically be an artificial intelligence model.
[0059] The framework layer 220 is used to initiate collective communication tasks based on the input data in the application layer 210. Specifically, the framework layer 220 can be an artificial intelligence model training framework, a software tool used to support the development, training, and deployment of artificial intelligence models. Examples include the PyTorch framework or the MindSpore framework.
[0060] The aggregate communication layer 230 is used to receive aggregate communication tasks sent by the framework layer 220, obtain the target data indicated by the target communication task, and store the target data in the communication buffer between the first computing cluster and the second computing cluster. When the amount of data in the communication buffer meets the preset conditions, the first computing cluster and the second computing cluster perform aggregate communication operations on the target data in the communication buffer according to the aggregate communication rules.
[0061] The driver layer 240 is used to invoke hardware resources to implement collection communication operations. These hardware resources include the GPU.
[0062] It should be noted that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0063] For ease of understanding, the collection communication method provided in this application is described below with reference to the accompanying drawings. This collection communication method is applicable to... Figure 3 The collective communication system 100 shown.
[0064] Figure 5 A flowchart illustrating a collection communication method provided in an embodiment of this application is shown. Figure 5 As shown, the collection communication method includes the following steps:
[0065] S101, the first computing cluster obtains the target data indicated by the target communication task.
[0066] In this embodiment of the application, the target data is the data to be transmitted as indicated by the target communication task, or the result data of at least one round of cluster-wide collective communication operations performed between computing nodes in the first computing cluster for the data to be transmitted.
[0067] For example, when the target communication task includes ReduceScatter, the target data is the result of at least one round of intra-cluster ensemble communication operations performed between computing nodes in the first computing cluster on the data to be transmitted. ReduceScatter is an operation in ensemble communication that combines reduction and scattering operations. ReduceScatter first performs a reduction operation on the input data on each computing node, and then scatters the reduced result back to each node, ensuring that each computing node ultimately receives only a portion of the reduced result.
[0068] In one possible implementation, if the target communication task includes ReduceScatter, a collective communication operation between computing nodes in the first computing cluster is performed to obtain the target data based on the data to be transmitted as indicated by the target communication task.
[0069] For example, before performing the collective communication operation between computing nodes in the first computing cluster, the first computing cluster stores the data to be transmitted indicated by the target communication task into the communication buffer according to the capacity of the communication buffer, so that the capacity utilization of the communication buffer reaches the maximum preset utilization rate.
[0070] For example, when updating model gradients using computing nodes in the first and second computing clusters, the ReduceScatter operation can reduce the gradients and distribute them among the computing nodes. First, the first computing cluster performs ReduceScatter operations between its internal computing nodes. The first computing cluster stores the data to be transmitted as indicated by the target communication task into the communication buffer according to its size and performs multiple rounds of ReduceScatter operations between its internal computing nodes. The result data obtained in each round of ReduceScatter operations includes the target data. Taking a target communication task indicating a data volume of 200MB and a communication buffer capacity of 50MB as an example, the first computing cluster divides the target communication task indicating a data volume of 50MB into four data slices. The first computing cluster then sequentially performs ReduceScatter operations between its internal computing nodes based on the data to be transmitted in each data slice.
[0071] In one possible implementation, where the target communication task includes AllGather, the target data is the data to be transmitted as indicated by the target communication task.
[0072] AllGather is a type of group communication operation. The AllGather operation means that each compute node in the computing cluster sends its data to all compute nodes in the cluster, so that each compute node ultimately possesses the data set from all compute nodes.
[0073] For example, in the case where the target communication task includes AllGather, the target data is the data to be transmitted as indicated by the target communication task. The first computing cluster determines the amount of target data; the amount of target data is a first preset threshold; the first preset threshold is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain.
[0074] For example, taking a communication buffer capacity of 50M, a target communication task indicating a data size of 50M, and a communication domain including a first computing cluster and a second computing cluster as an example, the first computing cluster determines a first preset threshold of 25M. Based on the first preset threshold, the first computing cluster divides the target communication task indicating a data size into two data slices, each with a data size of 25M. Therefore, the first computing cluster sequentially performs multiple rounds of aggregated communication operations with the second computing cluster for the target data corresponding to each data slice.
[0075] Optionally, if, during the process of segmenting the data to be transmitted indicated by the target communication task, the amount of data remaining in the target communication task is less than a data slice of the first preset threshold, then that portion of the data to be transmitted is still considered as a single data slice.
[0076] As can be seen from the above, when the first computing cluster and the second computing cluster perform the AllGather operation, the first computing cluster first performs an inter-cluster aggregate communication operation with the second computing cluster for the data to be transmitted as indicated by the target communication task. Therefore, the amount of the target data is the first preset threshold, which can ensure that the communication buffer is sufficient to store the result data of the AllGather operation between the first computing cluster and the second computing cluster.
[0077] S102, the first computing cluster stores the target data in the communication buffer between the first computing cluster and the second computing cluster.
[0078] The communication buffer is used to temporarily store data. The first computing cluster and the second computing cluster belong to the first communication domain, and the capacity of the communication buffer in the first communication domain is greater than the capacity of the communication buffer in the second communication domain. Computing nodes in the second communication domain belong to at most one computing cluster.
[0079] As shown above, compute nodes that need to perform aggregate communication across clusters require more cache resources to ensure communication latency and efficiency. Therefore, using differentiated configuration when configuring communication buffers can improve the utilization of communication buffers, thereby reducing latency during aggregate communication.
[0080] For example, before configuring the communication buffer, the computing cluster to which the computing nodes participating in the target communication task belong is identified. The capacity of the communication buffer corresponding to each computing node belonging to different computing clusters is greater than the capacity of the communication buffer corresponding to each computing node belonging to the same computing cluster.
[0081] Optionally, when creating a communication domain, the capacity of the communication buffer can be uploaded as a parameter. Specifically, the uploaded parameter size can vary depending on the number of computing clusters in the communication domain. The capacity of the communication buffer corresponding to at least two computing clusters in the communication domain is greater than the capacity of the communication buffer corresponding to a single computing cluster in the communication domain.
[0082] For example, by parsing a topology file or rank table, the computing cluster information to which each computing node participating in the target communication task belongs can be determined. Based on the computing cluster information to which each computing node belongs, the capacity of the communication buffer can be configured differently. The topology file includes the connection relationships between the computing nodes, specifically including the identifier and location of each computing node. The rank table is a collection of each computing node's number and related information.
[0083] S103, if the amount of data in the communication buffer meets the preset conditions, the first computing cluster and the second computing cluster perform a set communication operation on the target data in the communication buffer.
[0084] The preset conditions include the data capacity of the data to be transmitted in the communication buffer being a first preset threshold; or the data volume in the communication buffer not being increased within a preset time period.
[0085] For example, when the target communication task includes ReduceScatter, the target data is the result data of at least one round of cluster-wide aggregate communication operations performed between computing nodes in the first computing cluster for the data to be transmitted. First, the first computing cluster performs ReduceScatter operations between its individual computing nodes. The first computing cluster stores the data to be transmitted, as indicated by the target communication task, into a communication buffer according to the size of the communication buffer and performs multiple rounds of ReduceScatter operations between its individual computing nodes. The result data obtained in each round of ReduceScatter operations includes the target data. The first computing cluster stores the target data into the communication buffer according to the size of the communication buffer.
[0086] If the amount of data in the communication buffer does not increase within a preset time period, specifically, if the utilization rate of the communication buffer capacity reaches the maximum utilization rate, or if there is no result data of at least one round of cluster-based aggregate communication operation between computing nodes in the first computing cluster for the data to be transmitted within the preset time period; then, the first computing cluster and the second computing cluster perform ReduceScatter aggregate communication operation on the target data in the communication buffer to obtain the data to be output.
[0087] In other examples, where the target communication task includes an AllGather operation, the target data is the data to be transmitted as indicated by the target communication task. Before the first computing cluster stores the target data in the communication buffer between the first and second computing clusters, the first computing cluster determines the size of the target data. The size of the target data is a first preset threshold, which is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain. Based on the first preset threshold and the size of the data to be transmitted as indicated by the target communication task, the first computing cluster segments the data to be transmitted as indicated by the target communication task, obtaining at least one data slice, the size of which is the first preset threshold. At this point, the target data includes one data slice, and the first computing cluster stores one data slice in the communication buffer. When the size of the data to be transmitted in the communication buffer is the first preset threshold, the first and second computing clusters perform an AllGather collection communication operation on the target data in the communication buffer.
[0088] Optionally, if the data to be transmitted indicated by the target communication task is less than a first preset threshold, the first computing device stores the data to be transmitted indicated by the target communication task in the communication buffer, and the first computing cluster and the second computing cluster perform an AllGather collection communication operation on the target data in the communication buffer.
[0089] In one possible implementation, when the target communication task includes an AllGather operation, the first computing cluster stores the result data of the collective communication operation performed with the second computing cluster on the target data in the communication buffer into a communication buffer, such that the data volume of the communication buffer is a second preset threshold. Based on the result data of the collective communication operation in the communication buffer, the computing nodes in the first computing cluster perform collective communication operations. The second preset threshold is the ratio of the capacity of the communication buffer to the number of computing nodes in the first computing cluster.
[0090] For example, after the first computing cluster and the second computing cluster perform an AllGather set communication operation on the target data in the communication buffer, the first computing cluster stores the result data of the AllGather set communication operation with the second computing cluster on the target data in the communication buffer into the communication buffer. When the amount of data in the communication buffer is a second preset threshold, the first computing cluster performs at least one round of AllGather operation between each computing node in the cluster based on the result data of the set communication operation in the communication buffer to obtain the data to be output.
[0091] For example, if the number of computing nodes participating in the AllGather operation in the first computing cluster is 5 and the capacity of the communication buffer is 50M, and the amount of data of the result data of the set communication operation in the communication buffer is 10M, the first computing cluster performs the AllGather operation between each computing node based on the result data of the set communication operation in the communication buffer to obtain the final output data.
[0092] Optionally, if the amount of data resulting from the AllGather collective communication operation between the first computing cluster and the second computing cluster on the target data in the communication buffer is less than the second preset threshold, the result data of the AllGather collective communication operation between the first computing cluster and the second computing cluster on the target data in the communication buffer is stored in the communication buffer. Based on the result data of the collective communication operation in the communication buffer, the computing nodes in the first computing cluster perform collective communication operations.
[0093] As can be seen from the above, during the collective communication operation between the computing nodes of the first computing cluster and the second computing cluster, the utilization efficiency of the communication buffer usually affects the latency of the collective communication. Under the condition that the amount of data in the communication buffer meets the preset conditions, the first computing cluster and the second computing cluster perform collective communication operation on the target data in the communication buffer, which improves the utilization rate of the communication buffer, reduces the number of collective communication operations between the first computing cluster and the second computing cluster, and thus reduces the latency of the collective communication.
[0094] Taking the execution of the ReduceScatter operation in a scenario where a model is trained using the first computing cluster 110 and the second computing cluster 120 as an example, Figure 6 This diagram illustrates a communication process for a ReduceScatter operation according to an embodiment of this application. Figure 6 As shown, the first computing cluster 110 segments the data to be transmitted indicated by the ReduceScatter operation according to the size of the communication buffer 140, dividing the data to be transmitted into N data slices, including a first data slice, a second data slice, and an Nth data slice. The size of each data slice is equal to the size of the communication buffer 140. Here, N is a positive integer greater than 3.
[0095] First, the first computing cluster 110 stores the first data slice in the communication buffer 140. Each computing node in the first computing cluster 110 performs a ReduceScatter operation on the first data slice in the communication buffer 140. The result of the cluster-wide ReduceScatter operation performed between computing nodes in the first computing cluster 110 on the first data slice is the first target data. The first computing cluster 110 sequentially performs cluster-wide ReduceScatter operations on N data slices, generating N target data. The first computing cluster 110 temporarily stores the N target data in an input buffer. The input buffer and the communication buffer 140 are separate storage spaces.
[0096] Next, the first computing cluster 110 stores M target data into the communication buffer 140. Here, M is a positive integer less than or equal to N. If the amount of target data in the communication buffer 140 does not increase within a preset time period, the first computing cluster 110 and the second computing cluster 120 perform a ReduceScatter operation on the M target data in the communication buffer 140. The conditions under which the amount of target data in the communication buffer 140 does not increase within the preset time period include the communication buffer 140 being full, or the utilization rate of the communication buffer 140 being a preset maximum utilization rate, such as 100%. The conditions under which the amount of target data in the communication buffer 140 does not increase within the preset time period also include the first computing cluster not generating new target data within the preset time period.
[0097] Finally, the data to be output from the ReduceScatter operation of the M target data in the communication buffer 140 between the first computing cluster 110 and the second computing cluster 120 is temporarily stored in the output buffer, thus completing the ReduceScatter operation between the first computing cluster 110 and the second computing cluster.
[0098] In the above method, when the target communication task includes ReduceScatter, a hierarchical aggregate communication is adopted. First, ReduceScatter operations are performed within the cluster to temporarily store the target data within the cluster. Then, according to the capacity of the same communication buffer, cross-cluster ReduceScatter operations are performed on the target data to ensure the utilization rate of the communication buffer during the ReduceScatter operations between clusters, reduce cross-cluster ReduceScatter operations, and thus reduce the latency of aggregate communication.
[0099] Taking the AllGather operation as an example, in a scenario where the model is trained using the first computing cluster 110 and the second computing cluster 120, Figure 7 This diagram illustrates a communication process for an AllGather operation according to an embodiment of this application. Figure 7 As shown, the first computing cluster 110 segments the data to be transmitted indicated by the ReduceScatter operation according to the size of the communication buffer 140, dividing the data to be transmitted into N data slices, including a first data slice, a second data slice, and an Nth data slice. The data size of each data slice is a first preset threshold. Here, N is a positive integer greater than 3. The first preset threshold is the ratio of the capacity of the communication buffer 140 to the number of computing clusters in the communication domain.
[0100] First, during the AllGather operation, the target data is the data to be transmitted as indicated by the target communication task, specifically including a first data slice, a second data slice, and an Nth data slice. The first computing cluster 110 stores the first data slice in the communication buffer 140. The first computing cluster 110 performs an AllGather operation on the first data slice in the communication buffer 140 to obtain a first intermediate result. The first computing cluster 110 temporarily stores the first intermediate result in the output buffer, which stores the data to be output corresponding to the target communication task. The first computing cluster 110 sequentially performs an inter-cluster AllGather operation with the second computing cluster 120 for each data slice, and temporarily stores the intermediate result of each AllGather operation in the output buffer. If the data volume of the Nth data slice is less than a first preset threshold, the first computing cluster 110 directly stores the Nth data slice in the communication buffer 140, and the first computing cluster 110 performs an AllGather operation on the Nth data slice in the communication buffer 140.
[0101] Secondly, the first computing cluster 110 performs an AllGather operation among its internal computing nodes for each intermediate result. The first cluster 110 integrates and splits the intermediate results of each AllGather operation into M intermediate slices, including a first intermediate slice, a second intermediate slice, and an Mth intermediate slice. The data size of each intermediate slice is a second preset threshold, where M is a positive integer greater than 3. The second preset threshold is the ratio of the capacity of the communication buffer to the number of computing nodes in the first computing cluster 110. The first computing cluster 110 stores the first intermediate slice in the communication buffer 140, and each computing node in the first computing cluster 110 performs an AllGather operation on the first intermediate slice in the communication buffer 140.
[0102] Finally, the data to be output by each computing node in the first computing cluster 110 for performing AllGather operations on M intermediate slices is temporarily stored in the output cache, thus completing the AllGather operation between the first computing cluster 110 and the second computing cluster.
[0103] In the above method, when the target communication task includes AllGather, a hierarchical set communication approach is adopted. First, the target data is stored in the communication buffer according to a first preset threshold, and an AllGather operation is performed between sets. Then, an AllGather operation is performed between the computing nodes within the first computing cluster for the intermediate results. This not only ensures that the communication buffer is sufficient to store the result data generated by the AllGather operation, but also makes full use of the capacity of the communication buffer, ensuring the utilization rate of the communication buffer during the AllGather operation between clusters, reducing cross-cluster AllGather operations, and thus reducing the latency of set communication.
[0104] Optionally, in scenarios where the model is trained using the first computing cluster 110 and the second computing cluster 120, the AllGather operation is first executed to aggregate the data, followed by the AllGather operation for computation and distribution. This allows for updating model parameters, such as gradients. This avoids performing multiple rounds of AllReduce operations between the first computing cluster 110 and the second computing cluster 120, thus reducing latency during communication.
[0105] In summary, this application provides a collective communication method to improve the utilization of the communication buffer and further reduce latency during collective communication operations. This method is applied to a first computing cluster. The method includes: acquiring target data indicated by a target communication task; storing the target data in a communication buffer between the first and second computing clusters; and, when the amount of data in the communication buffer meets a preset condition, performing a collective communication operation with the second computing cluster on the target data in the communication buffer. As can be seen, storing the target data indicated by the target communication task in the communication buffer; and, when the amount of data in the communication buffer meets a preset condition, the first and second computing clusters perform a cross-computation cluster communication operation on the target data in the communication buffer. The second computing cluster is a computing cluster other than the first computing cluster. This method can improve the utilization of the communication buffer, reduce the number of cross-computation cluster collective communications, and thus effectively reduce link propagation latency during collective communication.
[0106] The above mainly describes the solution of the embodiments of this application from a methodological perspective. It is understood that the collection communication device, in order to achieve the above... Figure 5 The functions described herein include at least one of the hardware structures and software modules corresponding to the execution of each function. Those skilled in the art will readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.
[0107] This application embodiment can divide the first device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and is only a logical functional division. In actual implementation, there may be other division methods.
[0108] For example, Figure 8 A schematic diagram of the structure of a collective communication device provided in an embodiment of this application is shown. Figure 8 As shown, this aggregated communication device can be applied in a computing device, and the aggregated communication device 400 includes:
[0109] The acquisition module 410 is used to acquire the target data indicated by the target communication task;
[0110] The cache module 420 is used to store the target data in the communication buffer between the first computing cluster and the second computing cluster.
[0111] The communication module 430 is used to perform a collection communication operation with the second computing cluster on the target data in the communication buffer when the amount of data in the communication buffer meets the preset conditions.
[0112] In one possible implementation, the aggregation communication device 400 further includes a determination module for determining that, in the case that the target communication task includes AllGather, the target data is the data to be transmitted as indicated by the target communication task.
[0113] The data volume of the target data is determined; the data volume of the target data is a first preset threshold; the first preset threshold is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain.
[0114] In one possible implementation, the acquisition module 410 is further configured to perform a collective communication operation between computing nodes in the first computing cluster to acquire target data, based on the data to be transmitted indicated by the target communication task, if the target communication task includes ReduceScatter.
[0115] In one possible implementation, the cache module 420 is further configured to store the data to be transmitted indicated by the target communication task into the communication buffer according to the capacity of the communication buffer, so that the capacity utilization of the communication buffer reaches the maximum preset utilization rate.
[0116] In one possible implementation, the communication module 430 is further configured to, when the target communication task includes AllGather, store the result data of the collective communication operation with the second computing cluster on the target data in the communication buffer into the communication buffer, such that the data volume of the communication buffer is a second preset threshold; the second preset threshold is the ratio of the capacity of the communication buffer to the number of computing nodes in the first computing cluster; and perform collective communication operations between computing nodes in the first computing cluster based on the result data of the collective communication operation in the communication buffer.
[0117] The aforementioned collective communication device 400 can be applied to Figure 3 In the first computing cluster 110 of the collective communication system 100 shown.
[0118] The acquisition module 410, cache module 420, and communication module 430 can all be implemented in software or in hardware. For example, the implementation of the acquisition module 410 will be described below. Similarly, the implementation of the cache module 420 and communication module 430 can refer to the implementation of the acquisition module 410.
[0119] As an example of a software functional unit, the acquisition module 410 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the acquisition module 410 may include code running on multiple hosts / virtual machines / containers. It should be understood that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0120] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0121] As an example of a hardware functional unit, the acquisition module 410 may include at least one computing device, such as a server. Alternatively, the acquisition module 410 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0122] The multiple computing devices included in the acquisition module 410 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 410 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 410 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0123] It should be understood that in other embodiments, the acquisition module 410 can be used to execute any step in the collective communication method, and the cache module 420 and the communication module 430 can be used to execute any step in the operator testing method. The steps implemented by the acquisition module 410, the cache module 420 and the communication module 430 can be specified as needed. By implementing different steps in the collective communication method through the acquisition module 410, the cache module 420 and the communication module 430, all functions of the collective communication device 400 can be realized.
[0124] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform operations corresponding to any one of the implementation schemes and various feasible implementation schemes of the collective communication method.
[0125] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform operations corresponding to any one of the implementation schemes and various feasible implementation schemes of the set communication method.
[0126] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0127] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0128] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions, which, when loaded and executed on a computer, generate all or part of the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one network site, computer, server, or data center to another network site, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or it can be a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape, etc.), an optical medium (e.g., DVD, etc.), or a semiconductor medium (e.g., solid-state drive), etc.
[0129] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0130] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A collective communication method, characterized in that, Applied to a first computing cluster, the method includes: Acquire the target data indicated by the target communication task; The target data is stored in a communication buffer between the first computing cluster and the second computing cluster; When the amount of data in the communication buffer meets the preset conditions, a collective communication operation is performed with the second computing cluster for the target data in the communication buffer.
2. The method according to claim 1, characterized in that, The target data is the data to be transmitted as indicated by the target communication task, or the result data of at least one round of cluster-wide collective communication operations performed between computing nodes in the first computing cluster for the data to be transmitted.
3. The method according to claim 2, characterized in that, The preset conditions include the data capacity of the data to be transmitted in the communication buffer being a first preset threshold; or the data volume in the communication buffer not being increased within a preset time period; the first preset threshold is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain.
4. The method according to any one of claims 1-3, characterized in that, The first computing cluster and the second computing cluster belong to the first communication domain, and the capacity of the communication buffer is greater than the capacity of the communication buffer of the second communication domain; the computing nodes in the second communication domain belong to at most one computing cluster.
5. The method according to any one of claims 1-4, characterized in that, The acquisition of target data indicating the target communication task includes: If the target communication task includes ReduceScatter, the target data is obtained by performing a collective communication operation between computing nodes in the first computing cluster according to the data to be transmitted indicated by the target communication task.
6. The method according to claim 5, characterized in that, Before executing the aggregate communication operation between computing nodes in the first computing cluster based on the data to be transmitted indicated by the target communication task, the method further includes: Based on the capacity of the communication buffer, the data to be transmitted indicated by the target communication task is stored in the communication buffer, so that the capacity utilization of the communication buffer reaches the maximum preset utilization rate.
7. The method according to any one of claims 1-6, characterized in that, The acquisition of target data indicating the target communication task includes: In the case where the target communication task includes AllGather, the target data is the data to be transmitted as indicated by the target communication task; The data volume of the target data is determined; the data volume of the target data is the first preset threshold; the first preset threshold is the ratio of the capacity of the communication buffer to the number of computing clusters in the communication domain.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: In the case where the target communication task includes AllGather, the result data of the aggregate communication operation with the second computing cluster on the target data in the communication buffer is stored in the communication buffer, such that the data volume of the communication buffer is a second preset threshold; the second preset threshold is the ratio of the capacity of the communication buffer to the number of computing nodes in the first computing cluster; Based on the result data of the collective communication operation in the communication buffer, the computing nodes in the first computing cluster perform collective communication operations.
9. A collective communication device, characterized in that, The device includes: The acquisition module is used to acquire the target data indicated by the target communication task; A caching module is used to store the target data in a communication buffer between the first computing cluster and the second computing cluster; The communication module is used to perform a collective communication operation with the second computing cluster on the target data in the communication buffer when the amount of data in the communication buffer meets a preset condition.
10. A computing cluster, characterized in that, The computing cluster includes at least one computing device, which includes a processor, a memory, and computer programs / instructions stored in the memory; the processor executes the computer programs / instructions to enable the computing device to implement the method of aggregate communication as described in any one of claims 1-8.
11. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a computing device, cause the server to perform the method of aggregate communication as described in any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer program instructions that, when executed by a computing device, perform the method of collective communication as described in any one of claims 1-8.