A Cross-Cluster Data Processing Method and Apparatus

By splitting and placing data slicing between heterogeneous computing clusters, the problem of low communication efficiency between heterogeneous clusters is solved, and efficient cross-cluster data processing and large-scale training performance is improved.

CN120011112BActive Publication Date: 2025-06-20ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510488492.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-20
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The communication efficiency between heterogeneous computing clusters is inefficient, resulting in the overall efficiency of cross-cluster data processing.

Method used

By splitting the calculation result data into multiple data slices and performing standardized calculations in the host memory in sequence, efficient data transmission and processing between heterogeneous clusters are realized.

Benefits of technology

It significantly improves cross-cluster communication efficiency and data processing efficiency, shortens communication time, and optimizes the large-model training performance of heterogeneous clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011112B_ABST
    Figure CN120011112B_ABST
Patent Text Reader

Abstract

This specification discloses a cross-cluster data processing method and apparatus. The method includes: splitting the first result data stored in each computing node into multiple data slices; sequentially sending each data slice from each computing node to the host memory of the first computing cluster according to the order of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain the second result data corresponding to the received data slice; controlling the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; after each computing node receives the second result data corresponding to all data slices, obtaining the target calculation result. This solution improves the cross-cluster communication efficiency and further improves the cross-cluster data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular, to a cross-cluster data processing method and apparatus. Background Art

[0002] Collective communication, as a key communication mode, allows multiple computing nodes in a cluster to efficiently cooperate to complete complex computing tasks, and is widely applied in many fields such as distributed model training, data analysis, and scientific computing. Among them, during the collective communication process in a heterogeneous cluster, different computing clusters often consist of different types of computing nodes, such as Graphics Processing Unit (GPU), Central Processing Unit (CPU), Tensor Processing Unit (TPU), and dedicated AI chips. Different types or different manufacturers' computing nodes have different architectures and performance characteristics, so they show heterogeneity in hardware configuration.

[0003] Due to the hardware heterogeneity between heterogeneous computing clusters, it becomes particularly difficult to coordinate the communication between heterogeneous clusters. Especially in large-scale clusters, the gap between the computing nodes of different computing clusters leads to serious communication bottlenecks, affecting the communication efficiency between clusters.

[0004] Therefore, how to improve the communication efficiency between computing clusters and further improve the overall efficiency of cross-cluster data processing is an urgent problem to be solved. Summary of the Invention

[0005] This specification provides a cross-cluster data processing method and apparatus to partially solve the above problems existing in the prior art.

[0006] This specification adopts the following technical solutions:

[0007] This specification provides a cross-cluster data processing method, which is applied to a first computing cluster, and the first computing cluster is any one of multiple computing clusters that jointly execute a target task. The method includes:

[0008] Performing reduction calculation on the task data on each computing node in the first computing cluster to obtain first result data corresponding to the task data, and storing them separately on each computing node;

[0009] Splitting the first result data stored on each computing node into multiple data slices;

[0010] In the order of each data slice in the first result data, successively send each data slice from each computing node to the host memory of the first computing cluster, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster, to obtain the second result data corresponding to the received data slice, where the second computing cluster is any one of the multiple computing clusters other than the first computing cluster;

[0011] Control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; wherein, each computing node synchronously sends the data slice of the first data result stored by itself to the host memory;

[0012] After each computing node receives the second result data corresponding to all data slices, the target calculation result is obtained.

[0013] Optionally, successively sending each data slice from each computing node to the host memory of the first computing cluster specifically includes:

[0014] Successively send each data slice from each computing node to the host memory of the first computing cluster and store it in the first cache reservation area in the host memory of the first computing cluster;

[0015] Performing a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain the second result data corresponding to the received data slice specifically includes:

[0016] Perform a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain the second result data and store it in the second cache reservation area in the host memory of the first computing cluster, where the first cache reservation area and the second cache reservation area are isolated from each other.

[0017] Optionally, if there are multiple second computing clusters, the second result data obtained by the first computing cluster performing a reduction calculation on the received data slice and the data slices stored in the host memories of each second computing cluster are respectively stored in different second cache reservation areas.

[0018] Optionally, for each data slice, the data slice carries the address identifier of the computing node corresponding to the data slice;

[0019] Sending the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice specifically includes:

[0020] For each received data slice, according to the address identifier of the computing node corresponding to the received data slice, send the second result data corresponding to the received data slice from the host memory of the first computing cluster to the computing node corresponding to the received data slice.

[0021] Optionally, in the order of each data slice in the first result data, sequentially send each data slice from each computing node to the host memory of the first computing cluster, specifically including:

[0022] Start a process pool containing multiple data processing processes, where the number of the data processing processes is the same as the number of data slices obtained by splitting the first result data;

[0023] In the order of each data slice in the first result data, sequentially send each data slice from each computing node to the host memory of the first computing cluster through each processing process.

[0024] Optionally, there is a heterogeneous configuration between the computing nodes in the first computing cluster and the computing nodes in the second computing cluster.

[0025] Optionally, the target task includes a model training task.

[0026] This specification provides a data processing device for heterogeneous clusters, which is applied to a first computing cluster. The first computing cluster is any one of multiple computing clusters that jointly execute a target task, and includes:

[0027] A computing module, configured to perform reduction calculation on task data on each computing node in the first computing cluster, obtain first result data corresponding to the task data, and store them in each computing node respectively;

[0028] A splitting module, configured to split the first result data stored in each computing node into multiple data slices;

[0029] A first sending module, configured to sequentially send each data slice from each computing node to the host memory of the first computing cluster in the order of each data slice in the first result data, so that the host memory performs reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, where the second computing cluster is any one of the multiple computing clusters other than the first computing cluster;

[0030] A second sending module, configured to control the host memory to send the second result data from the host memory of the first computing cluster to the computing nodes corresponding to each received data slice while receiving subsequent data slices; wherein each of the computing nodes synchronously sends the data slices of the first data result stored by itself to the host memory.

[0031] A determination module, configured to obtain a target calculation result after each of the computing nodes receives the second result data corresponding to all the data slices.

[0032] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above cross-cluster data processing method.

[0033] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the above cross-cluster data processing method when executing the program.

[0034] At least one of the technical solutions adopted in this specification can achieve the following beneficial effects:

[0035] In the cross-cluster data processing method provided in this specification, the first result data stored by each computing node is split into multiple data slices; according to the order of each data slice in the first result data, each data slice is sequentially sent from each computing node to the host memory of the first computing cluster, so that the host memory performs a reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data corresponding to the received data slices; control the host memory to send the second result data from the host memory of the first computing cluster to the computing nodes corresponding to each received data slice while receiving subsequent data slices; after each computing node receives the second result data corresponding to all the data slices, obtain the target calculation result. This solution improves the cross-cluster communication efficiency and further improves the cross-cluster data processing efficiency.

[0036] As can be seen from the above method, since this solution splits the first data result into multiple data slices in advance before cross-cluster communication, and during the cross-cluster calculation process, the data slices are sent sequentially, the host memory can perform cross-cluster reduction calculation based on the received data slices first and return the second result data to the corresponding computing nodes in real time. In this way, the process of the computing nodes sending the first result data to the host memory and the process of the host memory sending the second result data to the computing nodes can be synchronized, shortening the communication time, significantly improving the communication efficiency between clusters, and further improving the overall efficiency of cross-cluster data processing. Description of the Drawings

[0037] The accompanying drawings described herein are used to provide a further understanding of the present specification, and constitute a part of the present specification. The illustrative embodiments of the present specification and their descriptions are used to explain the present specification, and do not constitute an improper limitation of the present specification. In the drawings:

[0038] Figure 1 It is a schematic flow chart of a cross-cluster data processing method provided in the present specification;

[0039] Figure 2 It is a schematic hardware architecture diagram of cross-cluster communication provided in the present specification;

[0040] Figure 3 It is an overall flow chart of cross-cluster communication provided in the present specification;

[0041] Figure 4 It is a schematic diagram of the time-consuming distribution of cross-cluster data processing provided in the present specification;

[0042] Figure 5 It is a transmission flow chart of slice data provided in the present specification;

[0043] Figure 6 It is an overall flow chart of the traditional cross-cluster communication method provided in the present specification;

[0044] Figure 7 It is a schematic diagram of a cross-cluster data processing device provided in the present specification;

[0045] Figure 8 It is provided in the present specification corresponding to Figure 1 schematic diagram of an electronic device. Specific Embodiments

[0046] To make the objectives, technical solutions, and advantages of the present specification clearer, the technical solutions of the present specification will be clearly and completely described below in conjunction with the specific embodiments of the present specification and the corresponding accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present specification.

[0047] In a heterogeneous cluster, the collective communication operation is one of the core components of distributed training. Through collective communication, the gradients on all computing nodes can be synchronized to ensure that each computing node has globally consistent model parameters. However, due to the large differences between different computing nodes in a heterogeneous cluster, traditional collective communication algorithms may not be able to fully utilize the advantages of each device, resulting in communication bottlenecks and low efficiency.

[0048] In recent years, a variety of optimization techniques have emerged to improve the efficiency of collective communication in heterogeneous clusters. For example, by analyzing the network topology of the cluster, dynamically select the optimal communication path and algorithm; use mixed precision for communication to reduce the amount of data transmission while maintaining training accuracy; through an asynchronous communication mechanism, enable computing and communication to proceed in parallel, reducing waiting time; for different levels of hardware, adopt a hierarchical communication strategy to utilize communication resources at different levels; through an intelligent scheduling algorithm, dynamically allocate tasks to different computing devices to ensure load balancing and avoid overloading some devices while leaving others idle.

[0049] Despite the significant progress made by these optimization techniques, there are still some major challenges in large model training for heterogeneous clusters. The complexity brought about by hardware heterogeneity makes it difficult to coordinate communication between different devices. Especially in large-scale clusters, the performance gap between devices may lead to serious load imbalance and communication bottlenecks. In addition, communication overhead accounts for a relatively large proportion in the entire data processing process. Especially during cross-cluster communication, data transmission takes a large amount of time, seriously affecting the overall efficiency.

[0050] Based on this, this specification provides a cross-cluster data processing method. By slicing and transmitting the calculation result data within the cluster, the data transmission process from device to host (D2H) and the data transmission process from host to device (H2D) can be synchronized, thereby minimizing communication overhead to the greatest extent under limited network bandwidth without affecting computing performance. The following combines the accompanying drawings to detail the technical solutions provided by each embodiment of this specification.

[0051] Figure 1 The following is a schematic flowchart of a cross-cluster data processing method provided in this specification, including the following steps:

[0052] S101: Perform reduction calculation on the task data on each computing node in the first computing cluster to obtain the first result data corresponding to the task data, and store them separately on each computing node.

[0053] In this specification, for each computing cluster, at least one computing node and at least one host memory can be set in the computing cluster. Each computing node is configured with a corresponding node memory (such as the video memory of a GPU). These computing nodes can include: GPU, CPU, NPU, TPU, neural network processing unit (NPU), and other AI chips. This specification does not make specific limitations on this.

[0054] Among them, the computing nodes of different computing clusters can be homogeneous nodes or heterogeneous nodes. Heterogeneous nodes can refer to different types of nodes (such as GPUs and NPUs), models of nodes, and manufacturers of nodes, resulting in a heterogeneous configuration among the computing nodes of the heterogeneous cluster. For the sake of understanding, this specification provides a schematic diagram of the hardware architecture for cross-cluster communication, such as Figure 2 shown.

[0055] Figure 2 is a schematic diagram of the hardware architecture for cross-cluster communication provided in this specification.

[0056] Among them, hostA and hostB are the hosts (host memories) of two heterogeneous computing clusters, and the computing nodes in each computing cluster are GPUs. The computing clusters communicate with each other through Remote Direct Memory Access (RDMA).

[0057] During the execution of tasks such as model training and big data processing, different computing nodes will obtain their corresponding execution results through calculations. At this time, it is necessary to perform data synchronization on the execution results corresponding to each computing node through collective communication operations to ensure global data consistency.

[0058] For example, during the distributed training of a business model, the computing nodes in each computing cluster can obtain the corresponding gradient data through calculations. At this time, it is necessary to synchronize the gradients on all computing nodes through collective communication to ensure that each computing node has globally consistent model parameters.

[0059] In this specification, the first computing cluster can pre-perform reduction calculations on the task data on each computing node it sets, obtain the first result data corresponding to each task data, and store them in the node memories of each computing node respectively. Among them, the first computing cluster is any one of multiple computing clusters jointly executing the target task.

[0060] For example, there are two computing nodes, GPU0 and GPU1, set in the first computing cluster. The task data on GPU0 is [[x1],[x2]], and the task data on GPU1 is [[x3],[x4]]. After reduction calculation, GPU0 obtains the first result data [x1 + x3] and stores it in its node memory, and GPU1 obtains the first result data [x2 + x4] and stores it in its node memory. For other computing clusters (hereinafter referred to as the second computing cluster), reduction calculation operations between computing nodes will also be performed.

[0061] Among them, the above-mentioned target task can be a model training task. Correspondingly, the task data can be the parameter data of the model training task (such as gradient data, feature vectors, loss values). Of course, the target task can also be other distributed computing tasks. Correspondingly, the task data can be transaction data, risk control data for distributed data processing tasks in the financial field, medical data for distributed data processing tasks in the medical field, etc.

[0062] S102: Split the first result data stored in each of the computing nodes into multiple data slices.

[0063] The first computing cluster can split the first result data stored in the memory of each of its nodes into multiple data slices.

[0064] For example, for the first result data stored in the node memory , it can be split into K data slices with a smaller granularity , that is , where K is an integer.

[0065] S103: According to the order of each data slice in the first result data, sequentially send each data slice from each of the computing nodes to the host memory of the first computing cluster, so that the host memory performs a reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data corresponding to the received data slices. Among them, the second computing cluster is any one of the multiple computing clusters other than the first computing cluster;

[0066] S104: Control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; among them, each of the computing nodes synchronously sends the data slices of its own stored first data result to the host memory.

[0067] The first computing cluster can sequentially send each data slice from the node memory of the computing node to the host memory of the first computing cluster according to the order of each data slice in the first result data.

[0068] For example, for the 9 data slices from GPU0 ~ , they can be arranged in the order from 1 to 9 in the first result data . The corresponding sending order of these data slices is 、 、 、……、 .

[0069] After receiving the data slices, the host memory of the first computing cluster can perform reduction calculations on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data corresponding to the received data slices.

[0070] After obtaining the second result data, the host memory can, while receiving subsequent data slices, send the second result data from the host memory of the first computing cluster to the computing nodes corresponding to each received data slice. Thus, the reception of the first result data and the transmission of the second result data are synchronized.

[0071] Continuing from the previous example, when the host memory of the first computing cluster receives after that, the host memory of the second computing cluster simultaneously receives At this time, the host memory of the first computing cluster performs a reduction calculation on and to obtain the second result data and immediately sends the second result data back to GPU0. In this way, while the first computing cluster performs a reduction calculation on and and returns to GPU0, the host memory of the first computing cluster is also concurrently receiving subsequent data slices The same is true for the second computing cluster.

[0072] To ensure that each second result data can be accurately sent to its corresponding computing node, for each data slice, the data slice can carry the address identifier of the computing node corresponding to the data slice. For each received data slice, the first computing cluster can, based on the address identifier of the computing node corresponding to the received data slice, send the second result data corresponding to the received data slice from the host memory of the first computing cluster to the computing node corresponding to the received data slice.

[0073] Among them, since the data slice carries its address identifier in the computing node, after performing a reduction calculation on the data slice, the second calculation result obtained by the host memory of the first computing cluster will still carry the address identifier. When the host memory sends the second result data to the computing node, it can determine which computing node the second result data needs to be sent to and the storage location in the node memory of the computing node based on the carried address identifier.

[0074] It should be noted that although the slice data in the first result data is serially sent to the host memory, the first result data in each computing node in the same computing cluster is synchronously sent to the host memory. For example, GPU0 sends the slice data ~ The process of sending to the host memory and the GPU1 synchronously perform the process of sending the sliced data ~ to the host memory.

[0075] For ease of understanding, this specification provides an overall flowchart of cross-cluster communication, as Figure 3 shown. Among them, the entire cross-cluster data processing flow includes the following steps:

[0076] Intra-cluster communication (reduce_scatter): Perform reduction calculations on [[x1],[x2]] on GPU0 and [[x3],[x4]] on GPU1, so as to obtain [[x1 + x3]] on GPU0 and [[x2 + x4]] on GPU1. Similarly, cluster 2 also performs the reduce_scatter operation, enabling each computing cluster to obtain the first result data and store it on the corresponding GPU.

[0077] Computing node - host memory data transfer (copy_d2h): Split into smaller-granularity data slices and copy them to the host memory.

[0078] Global collective communication (all_reduce): Perform reduction calculations between heterogeneous clusters. This process performs reduction operations on the data slices in the host memory of each computing cluster to obtain the address of the cache reservation area corresponding to a certain computing cluster, and perform reduction calculations on the data slices on the computing node and the corresponding data slices in the cache reservation areas of other computing clusters so that the corresponding data slices in the cache reservation areas of the two clusters and are updated with the same calculation results .

[0079] Host memory - computing node data transfer (copy_h2d): Synchronously update the reduction results of heterogeneous clusters. This step copies the reduction results synchronized in the host memory of each cluster back to the corresponding GPU to obtain the addresses of the cache reservation areas , , ……, and copy the final reduction calculation results one by one to , , ……, and then automatically copy them to the corresponding computing node Above, ensure that the calculation results of the entire diffusion reduction are synchronized on the corresponding computing nodes of each heterogeneous cluster.

[0080] Among them, the time-consuming distribution of cross-cluster data processing is as Figure 4 shown.

[0081] From Figure 4 it can be seen that through intelligent data sharding transmission and pipelined parallel execution of collective communication tasks, this solution realizes the synchronous progress of the h2d and d2h phases, significantly improving the data transmission efficiency.

[0082] In addition, this specification also provides the transmission process intention of the sliced data as Figure 5 shown, in order to facilitate the understanding of the data slice sending process.

[0083] As Figure 5 shown, during the copy_d2h process, the first result data is sequentially sent from the computing node (node memory) to the host memory in the form of data slices, while during the copy_h2d process, each second result data returns from the memory to the computing node where its corresponding data slice is located.

[0084] Furthermore, for step S102, after the first computing cluster sequentially sends each data slice from the computing node to the host memory of the first computing cluster, these data slices can be stored in the first cache reservation area in the host memory of the first computing cluster.

[0085] After the first computing cluster performs reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data, the second result data can be stored in the second cache reservation area in the host memory of the first computing cluster.

[0086] The data slices in the first cache reservation area need to participate in the cross-cluster reduction calculation and are isolated and stored in the second cache reservation area from the finally reduced second result data, ensuring the data correctness and the normal execution of the pipelined parallel tasks.

[0087] It should be added that for data processing scenarios with more than two computing clusters (such as multi-heterogeneous cluster distributed training scenarios), the processing process can trigger some reduction results to continue to perform reduction with the corresponding data slices in the cache reservation area of the third cluster and continue the diffusion reduction calculation according to the number of clusters.

[0088] Among them, the second result data obtained by the first computing cluster through reduction calculation on the received data slices and the data slices stored in the host memory of each second computing cluster are respectively stored in different second cache reservation areas, that is, multiple second cache reservation areas will be created in the host memory of each cluster. , , …….

[0089] In this specification, the above process can be implemented by a scheduling node (such as a CPU) in the computing cluster. The scheduling node can start a process pool containing multiple data processing processes, where the number of data processing processes is the same as the number of data slices obtained after splitting the first result data.

[0090] During the copy_d2h process, the scheduling node can send each data slice from the computing node to the host memory of the first computing cluster in sequence through each processing process according to the order of each data slice in the first result data.

[0091] In practical applications, a training framework of the model can be deployed in the scheduling node. The training framework splits an AllReduce operation into multiple subtasks corresponding to small-grained slice parameter data , so as to perform pipeline scheduling for each task. The parameter data of different slices need to perform the above serial operations and are executed in parallel in different processes to achieve overlapping between tasks (Overlapping), further improving the overall performance.

[0092] In the above process, the training framework can allocate subtasks , split the entire block of parameter data (the first result data) into small-grained parameter data slices , and copy them to the corresponding cache reservation areas , and then perform cross-cluster reduction calculation. The training framework automatically starts corresponding processes for each subtask , splitting the original single-process AllReduce operation into sub-processes with the corresponding slice number of K to process. The process pool of the training framework automatically binds the sub-processes generated by each subtask and obtains the corresponding parameter data slices , and executes the above steps in sequence. The training framework will automatically perform parallel pipelining on multiple sub-processes to form overlapping between tasks (Overlapping).

[0093] S105: After each computing node receives the second result data corresponding to all data slices, the target computing result is obtained.

[0094] For each computing cluster, when each computing node in the computing cluster has received the second result data of its respective corresponding all data slices, it indicates that the computing nodes in each computing cluster have completed data synchronization. At this time, the target computing result can be determined based on all the second result data stored in the node memory of each computing node, and subsequent tasks can be executed based on the target computing result.

[0095] For example, for a model training task, the target computing result can be the aggregated result of the gradient information obtained by all computing nodes. After the global synchronization of the gradient information is completed, each computing cluster can update the parameters of the model based on the target computing result.

[0096] In practical applications, the number of splits for data slices can be determined, so as to obtain the split number that maximizes the overall efficiency and split the result data during subsequent data processing.

[0097] Specifically, the time overhead for the process of copying data slices to the host memory can be denoted as Td2H, the time overhead for the reduction calculation process between different computing clusters can be denoted as Treduce, and the time overhead for copying the second result data to the computing nodes can be denoted as Th2d. The overall time overhead Tisum of this solution can be expressed as:

[0098]

[0099] And for Figure 6 the traditional cross-cluster communication method shown as follows, its overall time overhead Ttotal can be expressed as:

[0100]

[0101] From this, it can be obtained that the performance improvement ratio M of this solution compared to the traditional cross-cluster communication method can be expressed as:

[0102]

[0103] From the above formula, it can be seen that when M reaches the maximum value, the data slice size that contributes the most to the performance can be obtained, and then the number K ( ) of splits of the data slices can be obtained, which significantly improves the overall training performance of the heterogeneous cluster.

[0104] From the above method, it can be seen that this solution of the present invention significantly reduces the communication overhead of the AllReduce operation in the heterogeneous cluster through intelligent parameter sharding transmission and pipeline parallel processing, and greatly improves the overall performance of large model training.

[0105] Specifically, this method decomposes the original parameter data into smaller - granularity data slices, and through an optimized data transfer and reduction - calculation process, effectively reduces the data transfer time between the host memory and the node memory. At the same time, through multi - task parallel scheduling and pipeline overlapping technology, the utilization rate of computing resources is further improved, and the training cycle is shortened.

[0106] The experimental results show that, compared with the traditional training framework, this solution can significantly improve the efficiency of distributed training in a large - scale heterogeneous cluster environment. It not only optimizes the data transfer and reduction - calculation process, but also achieves significant performance improvement through an efficient parallel - processing mechanism, and is applicable to the distributed training scenario of large - scale heterogeneous clusters.

[0107] The above is a method for executing a computing task applied to collective communication in one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding device for executing a computing task applied to collective communication, as Figure 7 shown.

[0108] Figure 7 is a schematic diagram of a cross - cluster data - processing device provided by this specification, including:

[0109] A computing module 701, configured to perform reduction calculation on the task data on each computing node in the first computing cluster, obtain first result data corresponding to the task data, and store them in each of the computing nodes respectively;

[0110] A splitting module 702, configured to split the first result data stored in each computing node into multiple data slices;

[0111] A first sending module 703, configured to sequentially send each data slice from each computing node to the host memory of the first computing cluster according to the order of each data slice in the first result data, so that the host memory performs reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slices, where the second computing cluster is any computing cluster other than the first computing cluster among the multiple computing clusters;

[0112] A second sending module 704, configured to control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; wherein, each computing node synchronously sends the data slices of the first data result stored by itself to the host memory;

[0113] A determination module 705, configured to obtain a target calculation result after all the computing nodes receive the second result data corresponding to all the data slices.

[0114] Optionally, the first sending module 703 is specifically configured to sequentially send each data slice from each of the computing nodes to the host memory of the first computing cluster and store it in a first cache reservation area in the host memory of the first computing cluster.

[0115] Perform a reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data, and store it in a second cache reservation area in the host memory of the first computing cluster, where the first cache reservation area and the second cache reservation area are isolated from each other.

[0116] Optionally, if there are multiple second computing clusters, the second result data obtained by the first computing cluster performing a reduction calculation on the received data slices and the data slices stored in the host memory of each second computing cluster are respectively stored in different second cache reservation areas.

[0117] Optionally, for each data slice, the data slice carries an address identifier of the computing node corresponding to the data slice.

[0118] Optionally, the second sending module 704 is specifically configured to, for each received data slice, according to the address identifier of the computing node corresponding to the received data slice, send the second result data corresponding to the received data slice from the host memory of the first computing cluster to the computing node corresponding to the received data slice.

[0119] Optionally, the first sending module 703 is specifically configured to start a process pool including multiple data processing processes, where the number of data processing processes is the same as the number of data slices obtained by splitting the first result data; in the order of each data slice in the first result data, sequentially send each data slice from each of the computing nodes to the host memory of the first computing cluster through each processing process.

[0120] Optionally, the computing nodes in the first computing cluster and the computing nodes in the second computing cluster are in a heterogeneous configuration.

[0121] Optionally, the target task includes a model training task.

[0122] This specification also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 Provided a cross-cluster data processing method.

[0123] This specification also provides Figure 8 a schematic structural diagram of an electronic device corresponding to Figure 1 as shown. As Figure 8 described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned Figure 1 cross-cluster data processing method. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, but can also be hardware or logical devices.

[0124] The improvement of a technology can be clearly distinguished as either a hardware improvement (e.g., the improvement of circuit structures such as diodes, transistors, switches, etc.) or a software improvement (the improvement of method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method process into the hardware circuit. Therefore, it cannot be said that the improvement of a method process cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logic function is determined by the user's programming of the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog.

[0125] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0126] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0127] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0128] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0129] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0132] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0133] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0134] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0135] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0136] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0137] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0138] The various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0139] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A cross-cluster data processing method, characterized in that: The method is applied to a first computing cluster, where the first computing cluster is any one of a plurality of computing clusters that jointly execute a target task, and the method includes: Performing reduction calculation on the task data on each computing node in the first computing cluster to obtain first result data corresponding to the task data, and storing the first result data in each computing node respectively; Splitting the first result data stored in each computing node into multiple data slices; Sending each data slice from each computing node to the host memory of the first computing cluster in sequence according to the order of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; Controlling the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; After each computing node receives the second result data corresponding to all the data slices, the target computing result is obtained.

2. The method according to claim 1, characterized in that Sending each data slice from each computing node to the host memory of the first computing cluster in sequence specifically includes: Sending each data slice from each computing node to the host memory of the first computing cluster in sequence, and storing the data slice in a first cache reserved area in the host memory of the first computing cluster; Performing reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slices specifically includes: Perform reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data, and store them in the second cache reserved area in the host memory of the first computing cluster, wherein the first cache reserved area and the second cache reserved area are isolated from each other.

3. The method according to claim 2, characterized in that If there are multiple second computing clusters, the first computing cluster performs reduction calculations on the received data slices and the data slices stored in the host memory of each second computing cluster to obtain second result data, which are stored in different second cache reserved areas respectively.

4. The method according to claim 1, characterized in that For each data slice, the data slice carries the address identifier of the computing node corresponding to the data slice; Sending the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice specifically includes: For each received data slice, according to the address identifier of the computing node corresponding to the received data slice, the second result data corresponding to the received data slice is sent from the host memory of the first computing cluster to the computing node corresponding to the received data slice.

5. The method according to claim 1, characterized in that Sending the data slices from the computing nodes to the host memory of the first computing cluster in sequence according to the order of the data slices in the first result data specifically includes: Starting a process pool including a plurality of data processing processes, wherein the number of the data processing processes is the same as the number of data slices obtained after splitting the first result data; According to the order of each data slice in the first result data, each data slice is sent from each computing node to the host memory of the first computing cluster through each processing process in turn.

6. The method according to claim 1, characterized in that The computing nodes in the first computing cluster and the computing nodes in the second computing cluster are in a heterogeneous configuration.

7. The method according to claim 1, characterized in that The target task includes a model training task.

8. A heterogeneous cluster data processing device, characterized in that: Applied to a first computing cluster, where the first computing cluster is any one of a plurality of computing clusters that jointly execute a target task, including: A computing module, used for performing a reduction calculation on the task data on each computing node in the first computing cluster, obtaining first result data corresponding to the task data, and storing the first result data in each computing node respectively; A splitting module, used for splitting the first result data stored in each computing node into multiple data slices; A first sending module is used to send each data slice from each computing node to the host memory of the first computing cluster in sequence according to the sequence of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; A second sending module is used to control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving the subsequent data slice; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; The determination module is used to obtain the target calculation result after each computing node receives the second result data corresponding to all data slices.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Aggregate communication protocol calculation method and device, calculation card and storage medium

    CN119271617A

  • Data processing method and device

    CN119576844A