Cross-cluster data processing method and device

By performing standardized calculation and data slicing processing on task data in heterogeneous computing clusters, the problem of low communication efficiency between heterogeneous clusters is solved, and the efficiency of cross-cluster data processing is achieved.

CN120011112AActive Publication Date: 2025-05-16ZHEJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510488492.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The communication efficiency between heterogeneous computing clusters is inefficient, resulting in the overall efficiency of cross-cluster data processing.

Method used

By performing standardized calculations on the task data on the computing node, the first result data is obtained and split into multiple data slices. Then, in the order of data slices, they are sent to the host memory in turn, and cross-cluster standard calculations are performed, and the result data is finally synchronized to the computing node.

Benefits of technology

It improves cross-cluster communication efficiency, shortens communication time, and significantly improves the overall efficiency of cross-cluster data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011112A_ABST
    Figure CN120011112A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-cluster data processing method and device. The method comprises the following steps: splitting first result data stored in each computing node into a plurality of data slices; according to the sequence of each data slice in the first result data, the data slices are sent to a host memory of the first computing cluster from the computing nodes in sequence, so that the host memory carries out protocol computing on the received data slices and data slices stored in a host memory of a second computing cluster, obtaining second result data corresponding to the received data slices; when subsequent data slices are received in the control host, the second result data are sent to the computing nodes corresponding to the received data slices from the host memory of the first computing cluster; and after each computing node receives the second result data corresponding to all the data slices, obtaining a target computing result. According to the scheme, the cross-cluster communication efficiency is improved, and the cross-cluster data processing efficiency is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for cross-cluster data processing. Background Art

[0002] As a key communication mode, collective communication allows multiple computing nodes in a cluster to work together efficiently to complete complex computing tasks. It is widely used in many fields such as distributed model training, data analysis, scientific computing, etc. In the collective communication process of heterogeneous clusters, different computing clusters are often composed of different types of computing nodes, such as graphics processing units (GPUs), central processing units (CPUs), tensor processing units (TPUs), and dedicated AI chips. Computing nodes of different types or different manufacturers have different architectures and performance characteristics, so they are heterogeneous in hardware configuration.

[0003] Due to the heterogeneity of hardware between heterogeneous computing clusters, it is particularly difficult to coordinate communication between heterogeneous clusters. Especially in large-scale clusters, the gap between computing nodes in different computing clusters leads to serious communication bottlenecks, affecting the communication efficiency between clusters.

[0004] Therefore, how to improve the communication efficiency between computing clusters and further improve the overall efficiency of cross-cluster data processing is an urgent problem to be solved. Summary of the invention

[0005] This specification provides a cross-cluster data processing method and device to partially solve the above-mentioned problems existing in the prior art.

[0006] This manual adopts the following technical solutions: This specification provides a cross-cluster data processing method, the method is applied to a first computing cluster, the first computing cluster is any one of multiple computing clusters that jointly execute a target task, the method includes: Performing reduction calculation on the task data on each computing node in the first computing cluster to obtain first result data corresponding to the task data, and storing the first result data in each computing node respectively; Splitting the first result data stored in each computing node into multiple data slices; Sending each data slice from each computing node to the host memory of the first computing cluster in sequence according to the order of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; Controlling the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; After each computing node receives the second result data corresponding to all the data slices, the target computing result is obtained.

[0007] Optionally, sending the data slices from the computing nodes to the host memory of the first computing cluster in sequence specifically includes: Sending each data slice from each computing node to the host memory of the first computing cluster in sequence, and storing the data slice in a first cache reserved area in the host memory of the first computing cluster; Performing reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slices specifically includes: Perform reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data, and store them in the second cache reserved area in the host memory of the first computing cluster, wherein the first cache reserved area and the second cache reserved area are isolated from each other.

[0008] Optionally, if there are multiple second computing clusters, the first computing cluster performs reduction calculations on the received data slices and the data slices stored in the host memory of each second computing cluster and obtains second result data which are respectively stored in different second cache reserved areas.

[0009] Optionally, for each data slice, the data slice carries an address identifier of a computing node corresponding to the data slice; Sending the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice specifically includes: For each received data slice, according to the address identifier of the computing node corresponding to the received data slice, the second result data corresponding to the received data slice is sent from the host memory of the first computing cluster to the computing node corresponding to the received data slice.

[0010] Optionally, sending each data slice from each computing node to a host memory of the first computing cluster in sequence according to the sequence of each data slice in the first result data specifically includes: Starting a process pool including a plurality of data processing processes, wherein the number of the data processing processes is the same as the number of data slices obtained after splitting the first result data; According to the order of each data slice in the first result data, each data slice is sent from each computing node to the host memory of the first computing cluster through each processing process in turn.

[0011] Optionally, the computing nodes in the first computing cluster and the computing nodes in the second computing cluster are in a heterogeneous configuration.

[0012] Optionally, the target task includes a model training task.

[0013] This specification provides a heterogeneous cluster data processing device, which is applied to a first computing cluster, where the first computing cluster is any one of a plurality of computing clusters that jointly execute a target task, including: A computing module, used for performing a reduction calculation on the task data on each computing node in the first computing cluster, obtaining first result data corresponding to the task data, and storing the first result data in each computing node respectively; A splitting module, used for splitting the first result data stored in each computing node into multiple data slices; A first sending module is used to send each data slice from each computing node to the host memory of the first computing cluster in sequence according to the sequence of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; A second sending module is used to control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving the subsequent data slice; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; The determination module is used to obtain the target calculation result after each computing node receives the second result data corresponding to all data slices.

[0014] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned cross-cluster data processing method is implemented.

[0015] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned cross-cluster data processing method when executing the program.

[0016] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: In the cross-cluster data processing method provided in this specification, the first result data stored by each computing node is split into multiple data slices; each data slice is sent from each computing node to the host memory of the first computing cluster in sequence according to the order of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain the second result data corresponding to the received data slice; the host memory is controlled to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving the subsequent data slices; after each computing node receives the second result data corresponding to all the data slices, the target calculation result is obtained. This solution improves the efficiency of cross-cluster communication and further improves the efficiency of cross-cluster data processing.

[0017] It can be seen from the above method that since this solution will divide the first data result into multiple data slices before cross-cluster communication, the data slices are sent in sequence during the cross-cluster calculation process, so that the host memory can first perform cross-cluster reduction calculations based on the received data slices, and return the second result data to the corresponding computing node in real time. In this way, the process of the computing node sending the first result data to the host memory and the process of the host memory sending the second result data to the computing node can be carried out simultaneously, shortening the communication time, significantly improving the communication efficiency between clusters, and further improving the overall efficiency of cross-cluster data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings: Figure 1 A schematic diagram of a cross-cluster data processing method provided in this specification; Figure 2 A schematic diagram of a hardware architecture for cross-cluster communication provided in this specification; Figure 3 This is an overall flow chart of cross-cluster communication provided in this specification; Figure 4 This is a schematic diagram of the time consumption distribution of cross-cluster data processing provided in this manual; Figure 5 A transmission flow chart of slice data provided in this specification; Figure 6 This is an overall flow chart of the traditional cross-cluster communication method provided in this specification; Figure 7 A schematic diagram of a cross-cluster data processing device provided in this specification; Figure 8 A method corresponding to the Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0020] In heterogeneous clusters, collective communication operations are one of the core components of distributed training. Collective communication can synchronize the gradients on all computing nodes to ensure that each computing node has globally consistent model parameters. However, due to the large differences between different computing nodes in heterogeneous clusters, traditional collective communication algorithms may not be able to fully utilize the advantages of each device, resulting in communication bottlenecks and inefficiency.

[0021] In recent years, a variety of optimization techniques have emerged to improve the efficiency of collective communication in heterogeneous clusters. For example, by analyzing the network topology of the cluster, the optimal communication path and algorithm are dynamically selected; mixed precision is used for communication to reduce the amount of data transmission while maintaining training accuracy; asynchronous communication mechanisms are used to enable computing and communication to proceed in parallel to reduce waiting time; hierarchical communication strategies are adopted for different levels of hardware, and communication resources are utilized at different levels; and intelligent scheduling algorithms are used to dynamically assign tasks to different computing devices to ensure load balancing and avoid overloading some devices while idling others.

[0022] Despite the significant progress made in these optimization techniques, there are still some major challenges in training large models on heterogeneous clusters. The complexity brought by hardware heterogeneity makes it difficult to coordinate communication between different devices, especially in large-scale clusters, where the performance gap between devices can lead to serious load imbalance and communication bottlenecks. In addition, communication overhead accounts for a considerable proportion of the entire data processing process, especially when communicating across clusters, data transmission takes up a lot of time, which seriously affects the overall efficiency.

[0023] Based on this, this specification provides a cross-cluster data processing method, which uses a method for slicing and transmitting the calculation result data within the cluster so that the data transmission process from device to host (D2H) and the data transmission process from host to device (D2H) can be performed synchronously, thereby minimizing the communication overhead under limited network bandwidth without affecting the computing performance. The technical solutions provided by each embodiment of this specification are described in detail below in conjunction with the accompanying drawings.

[0024] Figure 1 This is a flow chart of a cross-cluster data processing method provided in this specification, which includes the following steps: S101: performing reduction calculation on task data on each computing node in the first computing cluster to obtain first result data corresponding to the task data, and storing the first result data in each computing node respectively.

[0025] In this specification, for each computing cluster, at least one computing node and at least one host memory can be set in the computing cluster, and each computing node is configured with corresponding node memory (such as GPU video memory). These computing nodes may include: GPU, CPU, NPU, TPU, neural network processor (Neural network Processing Unit, NPU) and other AI chips, which are not specifically limited in this specification.

[0026] The computing nodes of different computing clusters can be homogeneous nodes or heterogeneous nodes. Heterogeneous nodes can refer to different node types (such as GPU and NPU), node models, and node manufacturers, so that the computing nodes of heterogeneous clusters have heterogeneous configurations. For ease of understanding, this manual provides a hardware architecture diagram for cross-cluster communication, such as Figure 2 shown.

[0027] Figure 2 A schematic diagram of a hardware architecture for cross-cluster communication provided in this specification.

[0028] HostA and hostB are hosts (host memory) of two heterogeneous computing clusters, and the computing nodes in each computing cluster are GPUs. The computing clusters communicate with each other through Remote Direct Memory Access (RDMA).

[0029] When executing tasks such as model training and big data processing, different computing nodes will obtain their corresponding execution results through calculations. At this time, it is necessary to synchronize the execution results corresponding to each computing node through collective communication operations to ensure global consistency of the data.

[0030] For example, during the distributed training of a business model, each computing node in the computing cluster can obtain the corresponding gradient data through calculation. At this time, it is necessary to synchronize the gradients on all computing nodes through collective communication to ensure that each computing node has globally consistent model parameters.

[0031] In this specification, the first computing cluster can pre-calculate the task data on each computing node set therein, obtain the first result data corresponding to each task data, and store them in the node memory of each computing node respectively. The first computing cluster is any one of the multiple computing clusters that jointly execute the target task.

[0032] For example, the first computing cluster is provided with two computing nodes, GPU0 and GPU1, the task data on GPU0 is [[x1], [x2]], and the task data on GPU1 is [[x3], [x4]]. After the reduction calculation, GPU0 obtains the first result data [x1+x3] and stores it in its node memory, and GPU1 obtains the first result data [x2+x4] and stores it in its node memory. For other computing clusters (hereinafter referred to as the second computing cluster), the reduction calculation operation between computing nodes will also be performed.

[0033] Among them, the above-mentioned target task can be a model training task, and accordingly, its task data can be the parameter data of the model training task (such as gradient data, feature vectors, loss values). Of course, the target task can also be other distributed computing tasks, and accordingly, the task data can be transaction data and risk control data for distributed data processing tasks in the financial field, and medical data for distributed data processing tasks in the medical field, etc.

[0034] S102: Split the first result data stored in each computing node into multiple data slices.

[0035] The first computing cluster may split the first result data stored in the memory of each node therein into multiple data slices.

[0036] For example, for node memory The first result data stored in , which can be split into K data slices of smaller granularity ,Right now , where K is an integer.

[0037] S103: sending each data slice from each computing node to the host memory of the first computing cluster in sequence according to the sequence of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; S104: Control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory.

[0038] The first computing cluster may send the data slices from the node memory of the computing node to the host memory of the first computing cluster in sequence according to the sequence of the data slices in the first result data.

[0039] For example, for GPU0 ~ These 9 data slices can be found in the first result data In the order from 1 to 9, the sending order of these data slices is , , ,……, .

[0040] After receiving the data slice, the host memory can perform reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice.

[0041] After obtaining the second result data, the host memory can send the second result data from the host memory of the first computing cluster to the computing nodes corresponding to each received data slice while receiving the subsequent data slices, thereby synchronously receiving the first result data and sending the second result data.

[0042] Here we continue with the previous example. When the host memory of the first computing cluster receives After that, the host memory of the second computing cluster simultaneously receives , at this time, the host memory of the first computing cluster is and Execute the reduced calculation to obtain the second result data , and immediately the second result data Return to GPU0, so that in the first computing cluster and Execute the reduced calculation and return it to GPU0 At the same time, the host memory of the first computing cluster is also executing in parallel to receive subsequent data slices The same is true for the second computing cluster.

[0043] In order to ensure that each second result data can be accurately sent to its corresponding computing node, for each data slice, the data slice can carry the address identifier of the computing node corresponding to the data slice. For each received data slice, the first computing cluster can send the second result data corresponding to the received data slice from the host memory of the first computing cluster to the computing node corresponding to the received data slice according to the address identifier of the computing node corresponding to the received data slice.

[0044] Among them, since the data slice carries its address identifier in the computing node, after the data slice is reduced and calculated, the second calculation result obtained by the host memory of the first computing cluster will still carry the address identifier. When the host memory sends the second result data to the computing node, it can determine which computing node the second result data needs to be sent to and the storage location in the node memory of the computing node based on the address identifier it carries.

[0045] It should be noted that although the slice data in the first result data is sent serially to the host memory, the first result data in each computing node in the same computing cluster is sent synchronously to the host memory. For example, GPU0 sends the slice data ~ The process of sending to host memory and GPU1 will slice data ~ The sending to the host memory is done synchronously.

[0046] For ease of understanding, this specification provides an overall flow chart of cross-cluster communication, such as Figure 3 The entire cross-cluster data processing process includes the following steps: Intra-cluster communication (reduce_scatter): Reduce [[x1], [x2]] on GPU0 and [[x3], [x4]] on GPU1 to obtain [[x1+x3]] on GPU0 and [[x2+x4]] on GPU1. Similarly, cluster 2 also performs a reduce_scatter operation, so that each computing cluster obtains the first result data. , stored on the corresponding GPU.

[0047] Compute node-host memory data transfer (copy_d2h): Split into smaller data slices And copy to the host memory.

[0048] Global collective communication (all_reduce): Performs reduction calculations between heterogeneous clusters. This process reduces the data slices in the host memory of each computing cluster and obtains the cache reserved area corresponding to a computing cluster. Address, compute node Data slice on Cache reservation with other computing clusters The corresponding Data slice on Perform reduction calculations so that the two clusters correspond to the cache reservation area and Update the same calculation results .

[0049] Host memory-compute node data transfer (copy_h2d): Synchronous update of heterogeneous cluster protocol results. This step copies the host memory synchronization protocol results of each cluster back to the corresponding GPU to obtain the cache reserved area of ​​each heterogeneous cluster. , , ……, and copy the final calculation results to , , ..., and then automatically copied to the corresponding computing node This ensures that the entire diffusion reduction calculation results are synchronized on the corresponding computing nodes of each heterogeneous cluster.

[0050] The time consumption distribution of cross-cluster data processing is as follows: Figure 4 shown.

[0051] from Figure 4 It can be seen that this scheme realizes the synchronization of h2d and d2h stages through the pipeline parallel execution of intelligent data segmentation transmission and collective communication tasks, and significantly improves the data transmission efficiency.

[0052] In addition, this manual also provides Figure 5 The transmission flow of slice data shown is intended to facilitate understanding of the process of sending data slices.

[0053] like Figure 5 As shown, in the process of copy_d2h, the first result data is sent from the computing node (node ​​memory) to the host memory in sequence in the form of data slices, and in the process of copy_h2d, each second result data is returned from the memory to the computing node where its corresponding data slice is located.

[0054] Further, for step S102, after the first computing cluster sequentially sends each data slice from the computing node to the host memory of the first computing cluster, the data slices may be stored in the first cache reserved area in the host memory of the first computing cluster.

[0055] The first computing cluster performs reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain second result data, and then stores the second result data in the second cache reserved area in the host memory of the first computing cluster.

[0056] In the first cache reserved area The data slices need to participate in the cross-cluster reduction calculation and be stored in the second cache reserved area isolated from the second result data of the final calculation. , ensuring data correctness and normal execution of pipeline parallel tasks.

[0057] It should be added that for data processing scenarios with more than two computing clusters (such as multi-heterogeneous cluster distributed training scenarios), the processing process can trigger some of the reduction results to continue with the cache reservation area of ​​the third cluster. The correspondence in Data slice on Perform the reduction and continue the diffusion reduction calculation according to the number of clusters.

[0058] The first computing cluster performs a reduction calculation on the received data slices and the data slices stored in the host memory of each second computing cluster to obtain the second result data, which are stored in different second cache reserved areas, that is, multiple second cache reserved areas are opened in the host memory of each cluster. , , …….

[0059] In this specification, the above process can be implemented by a scheduling node (such as a CPU) in a computing cluster, and the scheduling node can start a process pool containing multiple data processing processes, where the number of data processing processes is the same as the number of data slices obtained after splitting the first result data.

[0060] During the copy_d2h process, the scheduling node may send each data slice from the computing node to the host memory of the first computing cluster through each processing process in sequence according to the sequence of each data slice in the first result data.

[0061] In actual applications, a model training framework can be deployed in the scheduling node. The training framework splits an AllReduce operation into multiple subtasks corresponding to small-grained slice parameter data. , so as to perform pipeline scheduling for each task. The parameter data of different slices need to undergo the above serial operations and be executed in parallel in different processes to achieve overlapping between tasks and further improve the overall performance.

[0062] In the above process, subtasks can be assigned by the training framework , split the entire parameter data (first result data) into small-grained parameter data slices , and copy it to the corresponding cache reserved area , and then perform cross-cluster reduction calculations. The training framework automatically calculates each subtask Start the corresponding process, split the original single process to perform AllReduce operation into sub-processes with the corresponding number of slices K for processing. The training framework process pool automatically binds the sub-process generated by each subtask and obtains the corresponding parameter data slices , execute the above steps in sequence, the training framework will automatically parallelize multiple sub-processes to form overlapping tasks.

[0063] S105: After each computing node receives the second result data corresponding to all the data slices, a target computing result is obtained.

[0064] For each computing cluster, when each computing node in the computing cluster receives the second result data of all corresponding data slices, it means that the computing nodes in each computing cluster have completed data synchronization. At this time, the target computing result can be determined according to all the second result data stored in the node memory of each computing node, and subsequent tasks can be performed based on the target computing result.

[0065] For example, for a model training task, the target calculation result can be the summary result of the gradient information obtained by all settlement nodes. After completing the global synchronization of the gradient information, each computing cluster can update the parameters of the model based on the target calculation result.

[0066] In practical applications, the number of splits of the data slices can be determined, so as to obtain the number of splits that maximizes the overall efficiency and split the result data for subsequent data processing.

[0067] Specifically, the time cost of copying data slices to host memory can be expressed as Td2H, the time cost of performing reduction calculations between different computing clusters can be expressed as Treducede, and the time cost of copying the second result data to the computing node can be expressed as Th2d. The overall time cost Tisum of this solution can be expressed as:

[0068] As for Figure 6 The overall time cost Ttotal of the traditional cross-cluster communication method shown in FIG. 1 can be expressed as:

[0069] It can be concluded that the performance improvement ratio M of this solution compared with the traditional cross-cluster communication method can be expressed as:

[0070] From the above formula, we can see that when M reaches its maximum value, we can get the data slice size that contributes most to performance, and then get the number of data slices split K ( ), which significantly improves the overall training performance of heterogeneous clusters.

[0071] It can be seen from the above method that this solution and the present invention significantly reduce the communication overhead of AllReduce operations in heterogeneous clusters through intelligent parameter sharding transmission and pipeline parallel processing, and greatly improve the overall performance of large model training.

[0072] Specifically, this method decomposes the original parameter data into smaller-granularity data slices, and effectively reduces the data transmission time between the host memory and the node memory through optimized data handling and simplified calculation processes. At the same time, through multi-task parallel scheduling and pipeline overlapping (Overlapping) technology, the utilization of computing resources is further improved and the training cycle is shortened.

[0073] Experimental results show that compared with traditional training frameworks, this solution can significantly improve the efficiency of distributed training in a large-scale heterogeneous cluster environment. It not only optimizes the process of data transfer and protocol calculation, but also achieves significant performance improvement through efficient parallel processing mechanisms. It is suitable for distributed training scenarios in large-scale heterogeneous clusters.

[0074] The above is one or more implementations of this specification applied to the computing task execution method of collective communication. Based on the same idea, this specification also provides a corresponding computing task execution device applied to collective communication, such as Figure 7 shown.

[0075] Figure 7 A schematic diagram of a cross-cluster data processing device provided for this specification includes: The computing module 701 is used to perform a reduction calculation on the task data on each computing node in the first computing cluster, obtain first result data corresponding to the task data, and store the first result data in each computing node respectively; A splitting module 702 is used to split the first result data stored in each computing node into multiple data slices; A first sending module 703 is used to send each data slice from each computing node to the host memory of the first computing cluster in sequence according to the sequence of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; A second sending module 704 is used to control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving the subsequent data slice; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; The determination module 705 is used to obtain the target calculation result after each computing node receives the second result data corresponding to all the data slices.

[0076] Optionally, the first sending module 703 is specifically used to send each data slice from each computing node to the host memory of the first computing cluster in sequence, and store it in a first cache reserved area in the host memory of the first computing cluster; Perform reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data, and store them in the second cache reserved area in the host memory of the first computing cluster, wherein the first cache reserved area and the second cache reserved area are isolated from each other.

[0077] Optionally, if there are multiple second computing clusters, the first computing cluster performs reduction calculations on the received data slices and the data slices stored in the host memory of each second computing cluster and obtains second result data which are respectively stored in different second cache reserved areas.

[0078] Optionally, for each data slice, the data slice carries an address identifier of a computing node corresponding to the data slice; Optionally, the second sending module 704 is specifically used to send, for each received data slice, the second result data corresponding to the received data slice from the host memory of the first computing cluster to the computing node corresponding to the received data slice according to the address identifier of the computing node corresponding to the received data slice.

[0079] Optionally, the first sending module 703 is specifically used to start a process pool including multiple data processing processes, wherein the number of the data processing processes is the same as the number of data slices obtained after splitting the first result data; and send each data slice from each computing node to the host memory of the first computing cluster through each processing process in turn according to the order of each data slice in the first result data.

[0080] Optionally, the computing nodes in the first computing cluster and the computing nodes in the second computing cluster are in a heterogeneous configuration.

[0081] Optionally, the target task includes a model training task.

[0082] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A cross-cluster data processing method is provided.

[0083] This manual also provides Figure 8 The one shown corresponds to Figure 1 A schematic diagram of the electronic device. Figure 8 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to the software implementation, this specification does not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0084] For the improvement of a technology, it can be clearly distinguished whether it is a hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented with "logic compiler" software, which is similar to the software compiler used when developing and writing programs. The original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog.

[0085] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0086] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0087] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0088] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0089] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0090] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0092] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0093] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0094] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0095] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0096] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0097] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0098] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0099] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A cross-cluster data processing method, characterized in that: The method is applied to a first computing cluster, where the first computing cluster is any one of a plurality of computing clusters that jointly execute a target task, and the method includes: Performing reduction calculation on the task data on each computing node in the first computing cluster to obtain first result data corresponding to the task data, and storing the first result data in each computing node respectively; Splitting the first result data stored in each computing node into multiple data slices; Sending each data slice from each computing node to the host memory of the first computing cluster in sequence according to the order of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; Controlling the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving subsequent data slices; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; After each computing node receives the second result data corresponding to all the data slices, the target computing result is obtained.

2. The method according to claim 1, characterized in that Sending each data slice from each computing node to the host memory of the first computing cluster in sequence specifically includes: Sending each data slice from each computing node to the host memory of the first computing cluster in sequence, and storing the data slice in a first cache reserved area in the host memory of the first computing cluster; Performing reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slices specifically includes: Perform reduction calculation on the received data slices and the data slices stored in the host memory of the second computing cluster to obtain the second result data, and store them in the second cache reserved area in the host memory of the first computing cluster, wherein the first cache reserved area and the second cache reserved area are isolated from each other.

3. The method according to claim 2, characterized in that If there are multiple second computing clusters, the first computing cluster performs reduction calculations on the received data slices and the data slices stored in the host memory of each second computing cluster to obtain second result data, which are stored in different second cache reserved areas respectively.

4. The method according to claim 1, characterized in that For each data slice, the data slice carries the address identifier of the computing node corresponding to the data slice; Sending the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice specifically includes: For each received data slice, according to the address identifier of the computing node corresponding to the received data slice, the second result data corresponding to the received data slice is sent from the host memory of the first computing cluster to the computing node corresponding to the received data slice.

5. The method according to claim 1, characterized in that Sending the data slices from the computing nodes to the host memory of the first computing cluster in sequence according to the order of the data slices in the first result data specifically includes: Starting a process pool including a plurality of data processing processes, wherein the number of the data processing processes is the same as the number of data slices obtained after splitting the first result data; According to the order of each data slice in the first result data, each data slice is sent from each computing node to the host memory of the first computing cluster through each processing process in turn.

6. The method according to claim 1, characterized in that The computing nodes in the first computing cluster and the computing nodes in the second computing cluster are in a heterogeneous configuration.

7. The method according to claim 1, characterized in that The target task includes a model training task.

8. A heterogeneous cluster data processing device, characterized in that: Applied to a first computing cluster, where the first computing cluster is any one of a plurality of computing clusters that jointly execute a target task, including: A computing module, used for performing a reduction calculation on the task data on each computing node in the first computing cluster, obtaining first result data corresponding to the task data, and storing the first result data in each computing node respectively; A splitting module, used for splitting the first result data stored in each computing node into multiple data slices; A first sending module is used to send each data slice from each computing node to the host memory of the first computing cluster in sequence according to the sequence of each data slice in the first result data, so that the host memory performs a reduction calculation on the received data slice and the data slice stored in the host memory of the second computing cluster to obtain second result data corresponding to the received data slice, wherein the second computing cluster is any computing cluster among the multiple computing clusters except the first computing cluster; A second sending module is used to control the host memory to send the second result data from the host memory of the first computing cluster to the computing node corresponding to each received data slice while receiving the subsequent data slice; wherein each computing node synchronously sends the data slice of the first data result stored by each computing node to the host memory; The determination module is used to obtain the target calculation result after each computing node receives the second result data corresponding to all data slices.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Aggregate communication protocol calculation method and device, calculation card and storage medium

    CN119271617A

  • Data processing method and device

    CN119576844A

  • Data interaction method and device between clusters, electronic equipment and storage medium

    CN119697196A

  • Data processing method, switching node, and related system

    WO2025002098A1