Aggregate communication method, computing device, network device and aggregate communication system

By introducing load reduction and broadcast requests between the computing device and the network device, combining the data buffering of the first temporary storage area and the second temporary storage area, the problems of low communication delay and resource utilization in large-scale multi-device scenarios under the Ring or Tree topology are solved, and efficient collective communication is achieved.

CN120342985APending Publication Date: 2025-07-18SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510686361.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing Ring or Tree communication topology causes a large number of data communication requests in large-scale multi-device scenarios, occupying network bandwidth, increasing communication delay, and computing devices have low resource utilization when waiting for communication, becoming a performance bottleneck in parallel computing systems.

Method used

The computing device sends load reduction requests and broadcast requests to the network device, and the network device uniformly processes data loading, reduction calculations and result broadcasts, reduces direct communication between computing devices, and uses the first temporary storage area and the second temporary storage area for data buffering to ensure efficient data transmission and processing between the computing device and the network device.

Benefits of technology

It reduces the computing task load and network traffic of computing devices, reduces communication delay, improves the resource utilization rate and collective communication efficiency of computing devices, avoids network congestion, and improves the overall computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342985A_ABST
    Figure CN120342985A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication, and provides an aggregate communication method, computing equipment, network equipment and an aggregate communication system.The method comprises the steps that a loading reduction request is sent to the network equipment to trigger the network equipment to load a target data block from a first temporary storage area of each computing equipment in a communication group, performing reduction operation on the basis of each loaded target data block; after receiving a reduction result returned by the network device, sending a broadcast request and the reduction result to the network device to trigger the network device to broadcast the reduction result to each computing device in the communication group; and storing the received reduction result broadcasted by the network equipment to the second temporary storage area, and writing the reduction result in the second temporary storage area to the output buffer area. According to the method, data loading, reduction calculation and result broadcasting are uniformly processed through the network equipment, the network communication traffic is reduced, the communication delay and the load of the calculation equipment are reduced, and the communication efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a collective communication method, a computing device, a network device, and a collective communication system. Background Art

[0002] In the fields of distributed computing and parallel computing, collective communication is crucial for ensuring data consistency and collaborative computing. Currently, the industry generally adopts Ring (ring-shaped) or Tree (tree-shaped) communication topologies to implement collective communication. However, although these two communication methods can achieve global data exchange, they both have obvious defects when dealing with large-scale multi-device scenarios.

[0003] In the above-mentioned Ring or Tree collective communication topologies, each computing device needs to exchange data with all other devices, which leads to a large number of data communication requests, occupies a large amount of network bandwidth, and frequent data transmission is prone to cause network congestion, significantly increasing communication latency. In addition, the device is idle while waiting for the communication to complete, resulting in waste of computing resources and a decrease in overall efficiency. These problems are particularly prominent when the scale of devices expands, becoming the main bottleneck restricting the performance of parallel computing systems. Summary of the Invention

[0004] The present invention provides a collective communication method, a computing device, a network device, and a collective communication system to solve the defects of large communication overhead, increased communication latency, and low utilization rate of computing resources existing in collective communication in related technologies.

[0005] The present invention provides a collective communication method, which is applied to a computing device, and the method includes: Segment the data in the input buffer and write each segmented data block into the first temporary storage area; Send a load reduction request to the network device, where the load reduction request is used to trigger the network device to load target data blocks from the first temporary storage area of each computing device in the communication group and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; After receiving the reduction result returned by the network device, send a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group; Store the reduction result broadcast by the network device received into the second temporary storage area, and write the reduction result in the second temporary storage area to the output buffer.

[0006] According to the collective communication method provided by the present invention, the step of writing each segmented data block into the first temporary storage area includes: When it is detected that the buffer status of the first temporary storage area is buffer ready, write each segmented data block into the first temporary storage area, and update the data status of the first temporary storage area to data ready; Correspondingly, the sending of the load reduction request to the network device includes: When it is detected that the data status of the first temporary storage area is data ready, send a load reduction request to the network device.

[0007] According to a collective communication method provided by the present invention, the sending of the broadcast request and the reduction result to the network device includes: When it is detected that the buffer status of the second temporary storage area is buffer ready, send a broadcast request and the reduction result to the network device.

[0008] According to a collective communication method provided by the present invention, after storing the reduction result broadcast by the network device received into the second temporary storage area, it further includes: Update the buffer status of the first temporary storage area to buffer ready, and update the data status of the second temporary storage area to data ready; Correspondingly, the writing of the reduction result in the second temporary storage area to the output buffer includes: When it is detected that the data status of the second temporary storage area is data ready, write the reduction result in the second temporary storage area to the output buffer, and update the buffer status of the second temporary storage area to buffer ready.

[0009] According to a collective communication method provided by the present invention, the number of loop times required for all computing devices in the communication group to execute tasks is determined based on the amount of data to be processed by each thread block and the amount of data processed by each computing device in one loop. The amount of data to be processed by each thread block is determined based on the total amount of data to be processed and the number of thread blocks. The amount of data processed by each computing device in one loop is determined based on the data block size and the number of devices in the communication group.

[0010] The present invention also provides a collective communication method, which is applied to a network device, and the method includes: Receive a load reduction request sent by any computing device, and based on the load reduction request, load target data blocks from the first temporary storage areas of each computing device in the communication group. The target data blocks are written into the first temporary storage areas from the input buffer after each computing device segments the data in the input buffer; Based on the loaded target data blocks, perform a reduction operation to obtain a reduction result, and return the reduction result to the any computing device, so that the any computing device sends a broadcast request and the reduction result to the network device; Receive the broadcast request and the reduction result sent by any one of the computing devices, and based on the broadcast request, broadcast the reduction result to each computing device within the communication group, so that each computing device stores the reduction result in the second scratchpad area and writes the reduction result in the second scratchpad area to the output buffer.

[0011] The present invention provides a collective communication device, which is applied to a computing device. The device includes: A data splitting unit, configured to split the data in the input buffer and write each split data block into the first scratchpad area; A reduction request unit, configured to send a load reduction request to a network device, where the load reduction request is used to trigger the network device to load target data blocks from the first scratchpad area of each computing device within the communication group and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; A broadcast request unit, configured to send a broadcast request and the reduction result to the network device after receiving the reduction result returned by the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device within the communication group; A result writing unit, configured to store the reduction result broadcast by the network device received into the second scratchpad area and write the reduction result in the second scratchpad area to the output buffer.

[0012] The present invention provides a collective communication device, which is applied to a network device. The device includes: A data loading unit, configured to receive a load reduction request sent by any one of the computing devices, and based on the load reduction request, load target data blocks from the first scratchpad area of each computing device within the communication group, where the target data blocks are split from the data in the input buffer by each computing device and written into the first scratchpad area from the input buffer; A reduction execution unit, configured to perform a reduction operation based on the loaded target data blocks to obtain a reduction result and return the reduction result to any one of the computing devices, so that any one of the computing devices sends a broadcast request and the reduction result to the network device; A result broadcast unit, configured to receive the broadcast request and the reduction result sent by any one of the computing devices, and based on the broadcast request, broadcast the reduction result to each computing device within the communication group, so that each computing device stores the reduction result in the second scratchpad area and writes the reduction result in the second scratchpad area to the output buffer.

[0013] The present invention also provides a computing device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, any one of the above-described collective communication methods applied to the computing device is implemented.

[0014] The present invention also provides a network device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, any one of the above-described collective communication methods applied to the network device is implemented.

[0015] The present invention also provides a collective communication system, including a plurality of the above-described computing devices and the above-described network device.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-described collective communication methods is implemented.

[0017] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, any one of the above-described collective communication methods is implemented.

[0018] For the collective communication method, computing device, network device, and collective communication system provided by the present invention, by sending a load reduction request and a broadcast request from the computing device to the network device, and having the network device uniformly process data loading, reduction calculation, and result broadcast, not only the computing tasks of the computing device itself are reduced, enabling the computing device to focus more on the core computing logic and improving the overall computing efficiency, but also frequent and large amounts of direct data exchange between computing devices are avoided, effectively reducing the network communication volume, reducing the communication links and data transmission paths, alleviating the network bandwidth pressure, thereby reducing the communication delay and improving the collective communication efficiency. By setting a first temporary storage area and a second temporary storage area on each computing device, not only can the efficient transfer and processing of data between the computing device and the network device be ensured, but also a certain buffering effect can be achieved to smooth the burst traffic of communication and avoid network congestion. Among them, the first temporary storage area is used to temporarily store each data block written from the input buffer, enabling the computing device to complete data segmentation and temporary storage locally without waiting for the cooperation of the network device or other computing devices, thereby avoiding the computing device from being idle due to waiting for communication to complete; similarly, the second temporary storage area allows the computing device to temporarily store the reduced result after receiving the broadcast, and then independently write the result to the output buffer later, avoiding the output buffer from being frequently occupied or blocked.

[0019] In addition, the network device broadcasts the reduction result to each computing device in the communication group only after receiving the broadcast request from the computing device. This design of actively triggering the broadcast operation by the computing device makes the broadcast operation process controllable, avoids resource conflicts, and reduces the complexity of the control logic of the network device at the same time. Brief Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 is a schematic structural diagram of the collective communication system provided by the present invention; Figure 2 is one of the schematic flowcharts of the collective communication method provided by the present invention; Figure 3 is a schematic structural diagram of the computing device provided by the present invention; Figure 4 is the second schematic flowchart of the collective communication method provided by the present invention; Figure 5 is a schematic architectural diagram of the method for realizing collective communication based on in-network computing provided by the present invention; Figure 6 is one of the schematic structural diagrams of the collective communication device provided by the present invention; Figure 7 is the second schematic structural diagram of the collective communication device provided by the present invention; Figure 8 is a schematic structural diagram of the device provided by the present invention. Detailed Embodiments

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0023] To facilitate understanding of the technical solutions and various embodiments provided by the present invention, the following first explains the key terms involved: Collective communication: It refers to communication and computing operations such as synchronization, data sending and receiving, etc. carried out by each distributed computing device in the fields of high-performance computing and distributed training, etc. parallel computing, mainly including Scatter, Reduce, Broadcast, AllReduce, Gather, AllGater, etc.

[0024] In-network computing: It means offloading computing operations from traditional computing devices to network devices to improve efficiency. This mode utilizes the computing power of network devices, enabling data to be processed during transmission, thereby reducing data transmission latency and improving overall computing efficiency.

[0025] Computing device: It refers to various hardware devices used to process data and execute computing tasks. For example, it can be a GPU (Graphics Processing Unit), GPGPU (General-purpose computing on Graphics Processing Units), TPU (Tensor Processing Unit), etc.

[0026] Network device: It refers to the dedicated hardware devices that make up an information communication network. For example, it can be a Switch, router, or other dedicated communication devices, etc.

[0027] In the fields of distributed computing and high-performance computing, collective communication in a multi-device (such as multi-GPU) environment is a key link to achieve efficient parallel computing. Collective communication operations, such as AllReduce, aim to integrate and compute data distributed on multiple computing devices and broadcast the results back to all devices, thereby ensuring data consistency among devices and supporting the smooth progress of tasks such as large-scale machine learning training and scientific computing.

[0028] To achieve this goal, communication topologies such as Ring or Tree are widely adopted in the industry. In the Ring structure, devices are connected in a ring, and data is gradually transmitted and reduced between adjacent devices in a fixed direction; the Tree structure constructs a tree-like hierarchy, where data converges and reduces from leaf nodes to the root node and then broadcasts back in the reverse direction. Both of these methods finally complete global collective communication through staged and batch-wise data exchange and computing.

[0029] However, these communication methods expose significant deficiencies when dealing with large-scale multi-device scenarios. First, each computing device needs to exchange data with all other devices, resulting in a large number of data communication requests. As the number of devices increases, the communication volume grows exponentially, severely occupying network bandwidth resources. Second, frequent data transmissions exacerbate network congestion. Especially when the data volume is large or the network bandwidth is limited, the communication delay increases significantly, becoming a bottleneck for the overall computing performance. In addition, the devices are idle while waiting for network communication to complete, unable to fully utilize computing resources, leading to low overall computing efficiency and extended task completion times.

[0030] The above deficiencies restrict the scalability and efficiency of multi-device parallel computing systems. Especially in fields such as deep learning that are sensitive to communication bandwidth and latency, there is an urgent need for more efficient collective communication strategies to optimize performance. In response, the present invention proposes a collective communication method. By sending load reduction requests and broadcast requests from computing devices to network devices, the network devices uniformly process data loading, reduction calculations, and result broadcasting, reducing the load and communication delay of computing devices, decreasing network communication volume, and improving the resource utilization rate of computing devices, thereby overcoming the above deficiencies.

[0031] Figure 1 is a schematic structural diagram of the collective communication system provided by the present invention, as Figure 1 shown. The system includes a network device 110 and multiple computing devices 120. These computing devices 120 are all connected to the network device 110, forming a communication network for collaborative work. Among them, the network device 110 (such as a switch) serves as the data exchange and processing center, responsible for receiving communication requests sent by the computing devices 120 (such as GPUs), coordinating data transmission between devices, performing core operations such as reduction calculations, and broadcasting the reduction results to all devices within the communication group to achieve efficient data distribution and processing. Here, in the collective communication system, when multiple computing devices need to collaborate to complete a computing task, these computing devices will be organized into a "communication group".

[0032] The computing device 120 undertakes parallel computing tasks, generates data to be processed and stores it in the local buffer. By sending requests (such as load reduction, broadcast) to the network device 110, it drives the flow and calculation of data in the network, and finally receives the processing results and completes subsequent tasks.

[0033] Specifically, the workflow of the collective communication system is as follows: Each computing device 120 (such as computing device 0) splits the data in its input buffer, temporarily stores it in the local first temporary storage area, and initiates a load reduction request to the network device 110. The network device 110 responds to the request, loads data blocks from each computing device within the same communication group, completes the reduction calculation, and returns the result to the computing device 120 (such as computing device 0) that initiated the request. This computing device 120 then initiates a broadcast request, and the network device 110 broadcasts the result to all computing devices within the communication group. Each device first temporarily stores the result in the local second temporary storage area and then writes it to the output buffer, thus completing a collective communication process.

[0034] In the embodiments of the present invention, by uniformly managing the reduction calculation and result broadcast by the network device, direct communication between computing devices can be reduced, the network complexity can be lowered. At the same time, the computing devices focus on the computing tasks, and the network device is responsible for data transmission and processing, realizing an efficient division of labor for computing and communication resources. Below, taking the computing device as a GPU and the network device as a switch as an example, the workflow of the collective communication system and the collective communication method provided by the present invention will be specifically introduced.

[0035] It should be noted that in the multi-GPU collective communication scenario, in order to efficiently utilize the parallel computing capabilities of multiple GPUs, large-scale computing tasks will be split. The purpose of task splitting is to process large-scale data in chunks and let different GPUs process these data chunks in parallel, thereby improving the overall computing efficiency. To facilitate the subsequent understanding of the collective communication method provided by the present invention, the task splitting will be introduced first.

[0036] Based on the above embodiments, the number of loops required for all computing devices within the communication group to execute the task is determined based on the amount of data to be processed by each thread block and the amount of data processed by each computing device in one loop. The amount of data to be processed by each thread block is determined based on the total amount of data to be processed and the number of thread blocks. The amount of data processed by each computing device in one loop is determined based on the data block size and the number of devices within the communication group.

[0037] Specifically, when splitting large-scale computing tasks, it is necessary to determine the amount of data processed by the computing device in each loop and the number of loops required to complete the entire task, so as to facilitate subsequent task allocation and scheduling in multi-GPU parallel computing and provide a basis for efficient multi-GPU collective communication and computing. Here, the number of loops represents the number of loops required to complete all data processing, that is, how many loops the entire task is split into for execution. The number of loops is calculated based on the amount of data to be processed by each thread block and the amount of data processed by each computing device in one loop.

[0038] It is understandable that in GPU programming, a thread block is one of the basic units for parallel execution on the GPU. Multiple thread blocks can execute different computational tasks in parallel. The amount of data to be processed by each thread block refers to the amount of data allocated to each thread block after splitting the task. Its calculation formula can be expressed as follows: Where, represents the amount of data to be processed by each thread block; represents the total amount of data to be processed (e.g., a tensor of size N); represents the number of thread blocks participating in the calculation, usually equal to the number of computing devices or a multiple thereof; represents the ceiling operation. It should be understood that through this calculation formula, the total amount of data to be processed can be evenly distributed among each thread block, avoiding uneven load. Since the total amount of data may not be divisible by the number of thread blocks, a ceiling operation is required to ensure that each thread block can at least process the allocated amount of data and guarantee that all data can be processed.

[0039] During the calculation process after task splitting, data is processed in batches (loops). After the data is chunked, each computing device will process a part of the data chunks in each loop. In each loop, the amount of data that a computing device needs to process is calculated based on the data chunk size and the number of devices in the communication group. Its calculation formula can be expressed as follows: In the above formula, represents the amount of data processed by each computing device in one loop; represents the size of a single data chunk, usually determined by the hardware characteristics of the computing device; represents the number of devices in the communication group.

[0040] After calculating and , the number of loops can be calculated through the following formula: In the above formula, represents the number of loops required to complete all data processing. Since there may be non - divisible cases in data amount allocation and calculation, a ceiling operation is required to ensure that all data can be processed and obtain the actual number of loops.

[0041] Exemplarily, taking the collective communication scenario of 4 GPUs (i.e., ) as an example, assume that these 4 GPUs need to cooperate to complete a task of size 1024 (i.e., The AllReduce operation of the tensors in () starts a thread block on each GPU (i.e., ), and the size of each data block is 64 (i.e., ). According to the above formula, it can be calculated that , which indicates that each thread block (i.e., each GPU) is allocated 256 elements; , that is, each computing device processes 256 elements in one loop; , indicating that all data processing can be completed in only 1 loop.

[0042] Specifically, the 1024 elements of the tensor can be divided into 4 blocks and allocated to 4 GPUs respectively, with 256 elements in each block. In the traditional AllReduce operation, each GPU needs to first reduce its own 64 elements and then synchronize the results through communication (such as Ring-AllReduce). For example, taking GPU0 as an example, it divides the 256 elements allocated to it into 4 blocks according to the number of GPUs, assuming they are data block 0, data block 1, data block 2, and data block 3, with 64 elements in each data block; the operations on GPU1, GPU2, and GPU3 are the same. Subsequently, GPU0 will communicate with GPU1, GPU2, and GPU3 respectively to obtain the corresponding data block 0 from each GPU, and then GPU0 performs a reduction calculation based on the local data block 0 and the data block 0 obtained from the other three GPUs (i.e., a total of 4 data blocks 0, a total of 256 elements). After obtaining the reduction result, it will communicate with GPU1, GPU2, and GPU3 again to transmit the reduction result to other GPUs. Similarly, GPU1, GPU2, and GPU3 will also perform the same operations as GPU0. It should be understood that although after data segmentation on each GPU, the obtained are all data block 0, data block 1, data block 2, and data block 3, the data block elements on different GPUs are not the same because they are segmented from different parts of the original tensor.

[0043] In the above collective communication process, each GPU needs to exchange data with all other GPUs, which will cause a significant increase in communication volume, seriously occupy network bandwidth, and significantly increase communication latency. In response, the present invention proposes a collective communication method based on in-network computing, which offloads operation tasks such as data loading, reduction calculation, and result broadcasting to network devices (such as switches), thereby enabling GPUs to avoid frequent data exchange and GPUs not to perform communication operations such as reduction and broadcasting, thus reducing network communication volume and reducing the load of computing devices and communication latency.

[0044] Based on any of the above embodiments, Figure 2 is one of the flow schematic diagrams of the collective communication method provided by the present invention, as Figure 2As shown, the method is applied to a computing device, and the method includes: Step 210: Split the data in the input buffer and write each split data block into the first temporary storage area.

[0045] It should be noted that the execution subject of the method provided in the embodiments of the present invention is a computing device, such as a GPU. To facilitate understanding of the technical solution provided in the embodiments of the present invention, the structure of the computing device will be briefly described first.

[0046] Figure 3 is a schematic structural diagram of the computing device provided by the present invention. As Figure 3 shown, each computing device 120 includes an input buffer 121 and an output buffer 122. Among them, the input buffer 121 refers to the area on the computing device for storing the original data to be processed. The data in the input buffer 121 is the original data to be subjected to a collective communication operation on each computing device. For example, in the above example of 4 GPUs collaborating to complete an AllReduce operation on a tensor data of size 1024, the data in the input buffer on each GPU is 256 elements allocated to each GPU. Before the collective communication starts, the data required for the computing task (such as reduction calculation) will be pre-loaded into the input buffer. These data may be read from an external storage device (such as a hard disk) or passed from other computing links.

[0047] The output buffer 122 refers to the area for storing the final result obtained after the collective communication process. When the collective communication operation (such as AllReduce) is completed, the computing device will store the finally obtained reduction result in the output buffer 122 for subsequent computing links to use or output to an external storage device.

[0048] In addition, to implement the collective communication method based on in-network computing, the embodiments of the present invention newly add a first temporary storage area 123 and a second temporary storage area 124 in the computing device. Among them, the first temporary storage area 123 is a temporary storage area on the computing device, mainly used for temporarily storing the split data blocks during the collective communication process. Before the computing device sends a load reduction request to the network device, the data in the input buffer will be split into multiple data blocks, and these data blocks will be written into the first temporary storage area so that the network device can conveniently read these data blocks for reduction operations. For example, in the above example of 4 GPUs collaborating to complete an AllReduce operation on a tensor data of size 1024, the first temporary storage area can store 4 data blocks (each block has 64 elements) obtained by further splitting the 256 elements allocated to each GPU. For example, GPU0 stores data block 0, data block 1, data block 2, and data block 3 in the first temporary storage area.

[0049] Similarly, the second temporary storage area is also a temporary storage area on the computing device, used to temporarily store the reduction results broadcast back from the network device. After each computing device receives the reduction results broadcast by the network device, it first stores the results in the second temporary storage area, and then writes the results from the second temporary storage area to the output buffer. It should be understood that the first temporary storage area and the second temporary storage area are two different temporary storage areas on the computing device. By separating the input and output temporary storage areas, the computing device can perform the operations of writing to the first temporary storage area and reading from the second temporary storage area simultaneously without waiting for global synchronization. The first temporary storage area is used to send the original data blocks, and the second temporary storage area is used to receive the reduction results. This separation avoids the competition problem when the same buffer is used for both sending and receiving.

[0050] Specifically, for each computing device, first, the data in the input buffer is sliced according to the number of devices in the communication group (such as 4 GPUs). For example, when 4 GPUs cooperate to complete an AllReduce operation on a tensor data of size 1024, the 1024 tensor elements are first evenly divided into 4 blocks, each block containing 256 elements, and are respectively assigned to 4 GPUs; then, each GPU further slices the 256 elements it is assigned into 4 blocks according to the number of devices, and each block contains 64 elements. Taking GPU0 as an example, it slices the 256 elements it is assigned into data block 0, data block 1, data block 2, and data block 3, each block containing 64 elements.

[0051] The sliced data blocks refer to the small pieces of data obtained after the original data in the input buffer undergoes a slicing operation. Each data block contains a certain number of elements, and these data blocks will serve as the basic units for subsequent collective communication operations.

[0052] After obtaining the sliced data blocks, the computing device can, through the internal storage management mechanism, unicast these data blocks from the input buffer to the first temporary storage area. For example, in a GPU, this usually involves operations on the video memory, and the data movement can be achieved through specific instructions or APIs (Application Programming Interfaces). Taking GPU0 as an example, it copies the sliced data block 0, data block 1, data block 2, and data block 3 from the input buffer to the first temporary storage area so that when the network device sends a load reduction request later, the network device can read these data blocks from the first temporary storage area. It should be understood that the above unicast method means that the operations of each computing device are independent and do not involve other computing devices. The data writing is point-to-point (from the local input buffer of the computing device to the local first temporary storage area), so it is unicast.

[0053] It is understandable that the first temporary storage area is used to temporarily store the data blocks written from the input buffer. Its existence enables the computing device to complete data segmentation and preparation locally without the need to immediately communicate with the network device. If the data in the input buffer is segmented and directly sent to the network device, the computing device needs to wait for the response of the network device to continue subsequent calculations, thereby reducing parallelism. This decoupling can prevent the computing device from being idle due to waiting for communication to complete, thus reducing the impact of communication latency on calculations.

[0054] Step 220: Send a load reduction request to the network device. The load reduction request is used to trigger the network device to load target data blocks from the first temporary storage area of each computing device within the communication group and perform a reduction operation based on the loaded target data blocks to obtain a reduction result.

[0055] Specifically, after the computing device writes the segmented data blocks into the first temporary storage area, it can send a load reduction request to the network device. Here, the computing device can send a load reduction request to the network device through a specific communication protocol and interface. For example, between a GPU and a switch, a dedicated communication library or driver can be set up, which provides functions or methods for sending requests. The GPU will call these functions, encapsulate the request information into a specific data packet, and then send the data packet to the switch through a network interface such as a PCIe (peripheral component interconnect express, high-speed serial computer expansion bus standard) bus.

[0056] The above load reduction request refers to an instruction sent by the computing device to the network device, which is used to inform the network device that specific data blocks need to be loaded and a reduction operation needs to be performed on these data blocks. It contains key information required for the network device to perform operations, such as device identifiers, data block identifiers, etc., so that the network device can accurately obtain the data and process it. Here, the device identifier refers to the identifier used to uniquely identify each computing device within the communication group. For example, in a multi-GPU collective communication scenario, each GPU has a unique identifier, such as GPU0, GPU1, GPU2, and GPU3, etc. The device identifier can help the network device determine the source computing device of the request so that the reduction result can be returned to the correct device later. The data block identifier refers to the identifier used to specify the data blocks that need to be loaded and reduced. For example, each GPU further divides the allocated data into multiple data blocks, and the data block identifier is used to identify these data blocks, such as data block 0, data block 1, etc. The network device can accurately read the corresponding data blocks from the first temporary storage area of the computing device according to the data block identifier.

[0057] Specifically, after receiving a load reduction request, the network device will parse the data block identifier carried in the request. Then, the network device will send an instruction to read the target data block to each computing device in the communication group through a specific communication mechanism. For example, the switch can send an instruction to all GPUs via multicast, requesting that the first scratchpad area of each GPU send the specified target data block to the switch. After receiving the instruction, the first scratchpad area of each GPU will send the data of the target data block to the switch through the network interface. It should be understood that the switch sending the read instruction via multicast means that the switch can send a request to all GPUs at once to read data blocks with the same identifier (such as data block 0), thereby reducing the communication overhead of the switch and improving efficiency. All GPUs can respond independently and send the data block to the switch.

[0058] It can be understood that the target data block is the data block that is currently specified to undergo a reduction operation among the data blocks split by the computing device. For example, when GPU0 sends a load reduction request to the switch, the target data block can be data block 0, and the switch will read data block 0 from the first scratchpad area of each GPU.

[0059] After receiving the target data blocks from each computing device, the network device will perform calculations on these data blocks according to a predetermined reduction operation to obtain a reduction result. For example, if the reduction operation is cumulative summation, the network device will add the elements at the corresponding positions in all target data blocks to obtain the final reduction result. It should be understood that the reduction operation is an operation that performs a specific calculation on multiple data. Common reduction operations include cumulative summation, finding the maximum value, finding the minimum value, etc. The reduction result refers to the result obtained after the reduction operation is executed. For example, in the case of cumulative summation, the reduction result is the sum of the elements at the corresponding positions in all target data blocks. The reduction result contains the comprehensive information obtained after the reduction operation and will be used for subsequent calculations or processing.

[0060] After the network device completes the reduction operation and obtains the reduction result, it will determine which computing device the reduction result needs to be returned to according to the device identifier obtained when parsing the request. Then, the network device will send the reduction result to the corresponding computing device through a specific communication mechanism. After receiving the reduction result, the computing device can store it in a register. For example, the computing device can store the reduction result in the TLR (Thread Local Register). It should be understood that the thread local register is a register in a computing device (such as a GPU) used to temporarily store thread-related data, and it has the characteristic of fast access. Storing the reduction result in the thread local register can facilitate subsequent threads to directly access and use these results. In GPU programming, threads are the basic units of parallel execution, and the thread local register can provide fast data access for threads, reduce the latency of memory access, and improve computing efficiency. Subsequent computing steps can directly read the reduction result from the thread local register for further processing.

[0061] Step 230, after receiving the reduction result returned by the network device, send a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group.

[0062] It should be noted that in collective communication (such as AllReduce), the reduction operation is only a part of the entire process. After completing the reduction operation, the reduction result needs to be broadcast to each computing device in the communication group to ensure that all computing devices have the same reduction result, thereby ensuring the correctness and consistency of subsequent calculations. For example, in distributed machine learning training, each GPU needs to know the global gradient reduction result to update its model parameters.

[0063] Specifically, after receiving the reduction result returned by the network device, the computing device can send a broadcast request and the reduction result to the network device through an interface provided by a specific communication library or driver. The computing device can encapsulate the reduction result in a specific data structure and send it to the network device through these interfaces (such as the PCIe bus) together with the broadcast request. Here, the broadcast request refers to an instruction sent by the computing device to the network device, used to inform the network device that specific data (i.e., the reduction result) needs to be broadcast to all computing devices in the communication group. It contains the data information of the reduction result and the target range of the broadcast (i.e., all computing devices in the communication group).

[0064] After receiving a broadcast request, the network device will parse the reduction result data and target range information in the request. Then, the network device will send the reduction result to each computing device in the communication group through a specific communication mechanism. For example, the network device can use multicast to send the data to each computing device. In the multicast mode, the network device only needs to send the data packet once, and the relevant devices in the network (i.e., each computing device in the communication group) will receive the data packet.

[0065] Step 240, store the reduction result broadcast by the received network device in the second temporary storage area, and write the reduction result in the second temporary storage area to the output buffer.

[0066] Specifically, when the network device broadcasts the reduction result to each computing device, the network interface of each computing device (such as the PCIe interface of the GPU) will receive this data. The communication library or driver of the computing device will parse the received data and store it in the pre-allocated second temporary storage area. Subsequently, the computing device will move the data in the second temporary storage area to the output buffer in a unicast manner through the internal storage management mechanism. For example, in the GPU, this can be achieved through video memory operations. The GPU will call specific instructions or APIs to copy the data in the second temporary storage area to the output buffer.

[0067] It can be understood that the second temporary storage area can be used as an intermediate buffer for temporarily storing the received data. Before writing the data to the output buffer, the computing device may need to perform some preprocessing operations on the data, such as data format conversion, verification, etc. The second temporary storage area provides space for these operations, avoiding the risks that may be brought by directly operating on the output buffer. For example, if the format of the received reduction result data is inconsistent with the format required by the output buffer, the computing device can perform format conversion on the data in the second temporary storage area and then write the converted data to the output buffer.

[0068] The method provided by the embodiments of the present invention sends a loading reduction request and a broadcast request from a computing device to a network device. The network device uniformly processes data loading, reduction calculation, and result broadcasting. This not only reduces the computing tasks of the computing device itself, enables the computing device to focus more on the core computing logic, and improves the overall computing efficiency, but also avoids frequent and large amounts of direct data exchange between computing devices, effectively reduces network traffic, reduces communication links and data transmission paths, alleviates network bandwidth pressure, thereby reducing communication latency and improving the efficiency of collective communication. By setting a first temporary storage area and a second temporary storage area on each computing device, it can not only ensure the efficient transfer and processing of data between the computing device and the network device, but also play a certain buffering role, smooth the burst traffic of communication, and avoid network congestion. Among them, the first temporary storage area is used to temporarily store each data block written from the input buffer, so that the computing device can complete data segmentation and temporary storage locally without waiting for the cooperation of the network device or other computing devices, thereby avoiding the computing device from being idle due to waiting for communication to complete; similarly, the second temporary storage area allows the computing device to temporarily store the reduction result after receiving the broadcast, and then independently write the result to the output buffer later, avoiding the output buffer from being frequently occupied or blocked. In addition, the network device will broadcast the reduction result to each computing device in the communication group only after receiving the broadcast request from the computing device. This design of actively triggering the broadcast operation by the computing device makes the broadcast operation process controllable, avoids resource conflicts, and at the same time reduces the complexity of the control logic of the network device.

[0069] Based on any of the above embodiments, in step 210, the writing of each segmented data block into the first temporary storage area includes: When it is detected that the buffer status of the first temporary storage area is buffer ready, write each segmented data block into the first temporary storage area and update the data status of the first temporary storage area to data ready.

[0070] It should be noted that the first temporary storage area is used to store the segmented data blocks, and these data blocks will participate in operations such as reduction later. Considering that when writing a data block into the first temporary storage area, if the space of the first temporary storage area is insufficient, the newly written data may overwrite the existing data or cause the writing to fail, resulting in data loss. In response to this, the embodiments of the present invention can ensure that the data writing operation is only performed when the buffer is ready by detecting the buffer status of the first temporary storage area, avoiding task interruption or exception caused by insufficient space.

[0071] Specifically, the buffer status of the first temporary storage area is used to indicate whether the temporary storage area is ready to receive new data. It reflects the current availability and writability of the temporary storage area. There are usually two states: buffer ready and buffer not ready. Buffer ready means that the temporary storage area has enough space to store new data blocks and can perform write operations; buffer not ready means that the temporary storage area may be occupied by other operations or has insufficient space and cannot receive new data temporarily.

[0072] When the computing device detects the buffer status of the first temporary storage area, it can be achieved through hardware flag bits, software status tracking, etc. For example, the hardware of the computing device can set specific flag bits or registers for each temporary storage area to record the buffer status. The computing device can judge the buffer status by reading the values of these flag bits or registers. For example, when the flag bit is "1", it means the buffer is ready, and when it is "0", it means the buffer is not ready. Another example is that in the operating system or driver of the computing device, a software status variable can be maintained to track the buffer status of the first temporary storage area. For example, set Head1 memory on the computing device to store the buffer status variable, and point to this memory through the HeadPtr1 pointer. The computing device reads the memory pointed to by the HeadPtr1 pointer in a polling manner to obtain this status variable and thus obtain the buffer status information.

[0073] When it is detected that the buffer status of the first temporary storage area is buffer ready, it indicates that the first temporary storage area is ready to receive the segmented data blocks. The computing device can write the data into this temporary storage area without worrying about problems such as insufficient space or other conflicts when writing the data.

[0074] After the computing device writes each segmented data block into the first temporary storage area, it can update the data status of the first temporary storage area to data ready. Here, the data status of the first temporary storage area is used to indicate whether valid data has been stored in this temporary storage area, that is, whether the data is ready to be used by subsequent operations (such as sending a load reduction request). There are also two states: data ready and data not ready. Data ready means that the segmented data blocks have been stored in the temporary storage area and subsequent operations such as sending reduction requests can be performed; data not ready means that there is no valid data in the temporary storage area or the data has not been completely written.

[0075] Specifically, after the computing device successfully writes each segmented data block into the first temporary storage area, the computing device can update the data status of the first temporary storage area to data ready through specific instructions or operations. For example, at the hardware level, the computing device can write a specific value to the register that controls the data status of the first temporary storage area to mark the data status as ready; at the software level, the driver can modify the corresponding status variable.

[0076] It is understandable that updating the data status of the first temporary storage area to data ready is to signal subsequent operations (such as sending a load reduction request), indicating that valid data has been stored in the first temporary storage area and relevant data processing and communication operations can be performed. This can ensure the orderliness and correctness of operations, avoiding subsequent operations being carried out when the data is not yet ready, resulting in errors or data inconsistencies.

[0077] Correspondingly, in step 220, the sending of the load reduction request to the network device includes: When it is detected that the data status of the first temporary storage area is data ready, sending a load reduction request to the network device.

[0078] Specifically, similar to detecting the buffer status, the computing device can detect the data status of the first temporary storage area by reading a hardware flag bit or querying a software status variable. The current data status information is stored in the hardware flag bit or the software status variable, and the computing device only needs to read this information to determine the data status. For example, on the computing device, Tail1 memory is set to store the data status variable, and the TailPtr1 pointer points to this memory. The computing device polls the memory pointed to by the TailPtr1 pointer to obtain this status variable, thereby obtaining the data status information.

[0079] When the computing device detects that the data status of the first temporary storage area is data ready, it indicates that the segmented data blocks have been stored in the first temporary storage area and these data are ready to be used by subsequent operations (such as sending a load reduction request). At this time, the computing device can safely send a load reduction request to the network device, requesting the network device to load the target data block in the first temporary storage area for reduction operations. In the embodiments of the present invention, by detecting the data status, the computing device can confirm whether the data is ready, avoiding operations such as sending a reduction request when the data is not yet ready, resulting in incorrect reduction results.

[0080] Based on any of the above embodiments, in step 230, the sending of the broadcast request and the reduction result to the network device includes: When it is detected that the buffer status of the second temporary storage area is buffer ready, sending the broadcast request and the reduction result to the network device.

[0081] It should be noted that, similar to the first temporary storage area, the second temporary storage area is used to store the reduction results broadcast by the network device. Before receiving the broadcast data, it is necessary to detect the buffer status to ensure that there is enough space in the temporary storage area to store this data. If the buffer status is buffer not ready, that is, the space is insufficient, continuously receiving broadcast data may cause data overflow, resulting in data loss or corruption. For example, in a large-scale distributed computing cluster, the reduction results may contain a large amount of data. If the buffer status is not detected, it may cause the second temporary storage area of some computing devices to overflow, affecting the correctness of the entire computing task.

[0082] Specifically, similar to detecting the buffer status of the first temporary storage area, the computing device can detect the buffer status of the second temporary storage area by means such as hardware flag bits and software status tracking. For example, a Head2 memory can be set on the computing device to store the buffer status variable of the second temporary storage area. The HeadPtr2 pointer points to this memory. The computing device reads the memory pointed to by the HeadPtr2 pointer in a polling manner to obtain this status variable, thereby obtaining the buffer status information of the second temporary storage area.

[0083] When the computing device detects that the buffer status of the second temporary storage area is buffer ready, it indicates that the second temporary storage area is ready to receive new data, that is, it has enough space to store the reduction results broadcast by the network device. The computing device can safely write the received reduction results into the second temporary storage area without worrying about insufficient space or data write conflicts.

[0084] Based on any of the above embodiments, in step 240, after storing the reduction results broadcast by the network device received into the second temporary storage area, it further includes: Updating the buffer status of the first temporary storage area to buffer ready and updating the data status of the second temporary storage area to data ready.

[0085] Specifically, when the computing device stores the reduction results broadcast by the network device into the second temporary storage area, it indicates that the network device has completed the reduction calculation of the data block in the first temporary storage area. In this case, the buffer status of the first temporary storage area can be updated to buffer ready, indicating that this temporary storage area has completed the previous data processing and is ready to receive new data again, so that the first temporary storage area can receive a new round of data blocks. At the same time, update the data status of the second temporary storage area to data ready, indicating that valid reduction result data has been stored in the second temporary storage area and subsequent data processing and transmission operations can be performed.

[0086] Specifically, when the computing device updates the buffer status of the first scratchpad area, it can update the buffer status by writing a specific value to the hardware register that controls the buffer status of the first scratchpad area. For example, set a certain flag bit in the register to a value indicating buffer readiness (such as "1"). It can also update the buffer status of the first scratchpad area by calling specific software instructions or functions, which will modify the status variables maintained in the software. For example, modify the buffer status variable in the Head1 memory pointed to by the HeadPtr1 pointer, so as to update the buffer status information.

[0087] Similar to updating the buffer status, the computing device can update the data status by writing a specific value to the hardware register that controls the data status of the second scratchpad area. For example, set a flag bit to indicate that the data is ready. Another example is that at the software level, the driver or application can update the data status of the second scratchpad area by modifying the corresponding software status variables.

[0088] Correspondingly, in step 240, writing the reduction result in the second scratchpad area to the output buffer includes: When it is detected that the data status of the second scratchpad area is data ready, write the reduction result in the second scratchpad area to the output buffer, and update the buffer status of the second scratchpad area to buffer ready.

[0089] Specifically, similar to detecting the buffer status, the computing device can detect the data status of the second scratchpad area by reading the hardware flag bit or querying the software status variable. The current data status information is stored in the hardware flag bit or software status variable, and the computing device only needs to read this information to judge the data status. For example, set the Tail2 memory on the computing device to store the data status variable of the second scratchpad area, and point to this memory through the TailPtr2 pointer. The computing device reads the memory pointed to by the TailPtr2 pointer in a polling manner to obtain this status variable, so as to obtain the data status information of the second scratchpad area.

[0090] When the computing device detects that the data status of the second scratchpad area is data ready, it indicates that valid reduction result data has been stored in the second scratchpad area, and this data is ready to be used by subsequent operations (such as writing to the output buffer). At this time, the computing device can safely write the reduction result in the second scratchpad area to the output buffer to ensure the correct transmission and processing of the data. After the computing device writes the reduction result of the second scratchpad area to the output buffer, it can update the buffer status of the second scratchpad area to buffer ready, so that the second scratchpad area can receive a new round of reduction results.

[0091] Based on any of the above embodiments, Figure 4 is the second flowchart of the collective communication method provided by the present invention, asFigure 4 As shown in Figure 4 , this method is applied to a network device, and the method includes: Step 410: Receive a load reduction request sent by any computing device. Based on the load reduction request, load target data blocks from the first scratchpad area of each computing device in the communication group. The target data blocks are obtained by each computing device slicing the data in the input buffer and then writing the sliced data blocks from the input buffer to the first scratchpad area. Step 420: Based on the loaded target data blocks, perform a reduction operation to obtain a reduction result, and return the reduction result to the any computing device, so that the any computing device sends a broadcast request and the reduction result to the network device. Step 430: Receive the broadcast request and the reduction result sent by the any computing device, and based on the broadcast request, broadcast the reduction result to each computing device in the communication group, so that each computing device stores the reduction result in the second scratchpad area and writes the reduction result in the second scratchpad area to the output buffer.

[0092] Specifically, the execution subject of the method provided in the embodiments of the present invention may be a network device, such as a switch. Before executing step 410, each computing device will first slice the data in the local input buffer, and then write the sliced data blocks from the input buffer to the first scratchpad area. Subsequently, each computing device will send a load reduction request to the network device, and the request will carry the corresponding device identifier and data block identifier.

[0093] After the network device receives the load reduction request of any one computing device, it will parse the request. According to the data block identifier obtained by parsing, it can determine the target data blocks that need to be loaded from the first scratchpad area of each computing device in the communication group. Then, the network device can use multicast to send a read instruction for the target data blocks to all computing devices. After each computing device receives the instruction, it will send the target data blocks in the first scratchpad area to the network device.

[0094] After the network device receives the target data blocks from each computing device, it will perform a reduction operation based on these target data blocks to obtain the corresponding reduction result. Subsequently, the network device will determine which computing device needs to return the reduction result according to the device identifier obtained when parsing the request, and send the reduction result to the corresponding computing device through a specific communication mechanism.

[0095] After receiving the reduction result, the computing device can store it in the thread-local register. Then, it sends a broadcast request and the reduction result to the network device through the interface provided by a specific communication library or driver. After receiving the broadcast request, the network device will parse the reduction result data and the target range information in the request. Then, the network device broadcasts the reduction result to each computing device in the communication group. After each computing device receives the reduction result broadcast by the network device, it first stores it in a pre-allocated second scratchpad area. Subsequently, the computing device moves the data in the second scratchpad area to the output buffer in a unicast manner through the internal storage management mechanism, thus completing a collective communication process.

[0096] The method provided by the embodiments of the present invention sends a load reduction request and a broadcast request from the computing device to the network device, and the network device uniformly processes data loading, reduction calculation, and result broadcasting. This not only reduces the computing tasks of the computing device itself, enables the computing device to focus more on the core computing logic, and improves the overall computing efficiency, but also avoids frequent and large direct data exchanges between computing devices, effectively reduces the network traffic, reduces the communication links and data transmission paths, alleviates the network bandwidth pressure, thereby reducing the communication delay and improving the collective communication efficiency. By setting a first scratchpad area and a second scratchpad area on each computing device, it can not only ensure the efficient transfer and processing of data between the computing device and the network device, but also play a certain buffering role, smooth the burst traffic of communication, and avoid network congestion. Among them, the first scratchpad area is used to temporarily store each data block written from the input buffer, so that the computing device can complete data segmentation and temporary storage locally without waiting for the cooperation of the network device or other computing devices, thereby avoiding the computing device from being idle due to waiting for communication to complete; similarly, the second scratchpad area allows the computing device to temporarily store the broadcast reduction result and then independently write the result to the output buffer later, avoiding the output buffer from being frequently occupied or blocked. In addition, the network device will broadcast the reduction result to each computing device in the communication group only after receiving the broadcast request from the computing device. This design of actively triggering the broadcast operation by the computing device makes the broadcast operation process controllable, avoids resource conflicts, and at the same time reduces the control logic complexity of the network device.

[0097] Based on any of the above embodiments, Figure 5 is a schematic architecture diagram of the method for realizing collective communication based on in-network computing provided by the present invention, as Figure 5As shown, in the embodiment of the present invention, the computing device is a GPU, and the network device is a switch. Taking the process of 4 GPUs (such as GPU0, GPU1, GPU2, and GPU3, where GPU2 and GPU3 are not shown in the figure) collaborating to complete an AllReduce operation on a tensor data of size 1024 as an example. First, the tensor data is divided into 4 blocks, with 256 elements in each block, and are respectively assigned to 4 GPUs. Figure 5 The data stored in the input buffers of GPU0 and GPU1 in Figure 5 are the 256 elements assigned to them. The AllReduce process is mainly divided into four stages: S1, Scatter: Locally on each GPU, the data in the input buffer is split according to the number of GPUs, and the split data blocks are written from the input buffer into the first scratchpad area. For example, taking GPU0 as an example, it splits the data in its local input buffer into 4 blocks, namely data block 0, data block 1, data block 2, and data block 3, with 64 elements in each block, and then these 4 data blocks are unicasted from the input buffer to the first scratchpad area.

[0098] It should be noted that assuming the data state in the input buffer is always data ready, before writing the 4 split data blocks into the first scratchpad area, GPU0 needs to first detect the buffer state of the first scratchpad area. For example, GPU0 can set a Head1 memory locally to store the buffer state of the first scratchpad area, point to this memory through the HeadPtr1 pointer, and read the buffer state in the memory pointed to by the HeadPtr1 pointer in a polling manner. Initially, the buffer state of the first scratchpad area is set to buffer ready. When GPU0 reads the buffer state as buffer ready through polling, it can move the 4 data blocks in the input buffer to the first scratchpad area in parallel through multiple warps.

[0099] Subsequently, update the buffer state of the first scratchpad area to buffer not ready to avoid subsequent abnormal or failed data writing to the first scratchpad area. At the same time, update the data state of the first scratchpad area to data ready. Here, GPU0 can set a Tail1 memory locally to store the data state of the first scratchpad area, point to this memory through the TailPtr1 pointer, and when updating the data state of the first scratchpad area, update the variable in the memory pointed to by the TailPtr1 pointer.

[0100] Similarly, GPU1, GPU2, and GPU3 will locally further divide the 256 elements allocated to them into 4 blocks, namely data block 0, data block 1, data block 2, and data block 3, each with 64 elements, and write these data blocks from their respective input buffers to their respective first temporary storage areas. After writing is completed, the buffer status and data status of their respective first temporary storage areas are updated.

[0101] S2, Reduce: When each GPU detects that the data status of the first temporary storage area is data ready, it will send a load reduction request to the switch. For example, taking GPU0 as an example, it can use polling to read the data status in the memory pointed to by the TailPtr1 pointer. When the data status is detected to be data ready, it indicates that the data in the first temporary storage area is ready. At this time, GPU0 can send a load reduction request to the switch, which carries a device identifier (that is, GPU0) and a data block identifier (such as data block 0).

[0102] After receiving the load reduction request from GPU0, the switch will parse the request to obtain the device identifier and data block identifier. Then the switch sends an instruction to read data block 0 to the first buffer of all GPUs (i.e., GPU0~GPU3) through multicast. After receiving the instruction, the first buffer of each GPU sends data block 0 to the switch.

[0103] After the switch receives data blocks 0 from all GPUs, it will perform reduction operations (such as cumulative summation) on these data blocks 0 to obtain the corresponding reduction results. Subsequently, the switch will return the reduction results to the TLR register of GPU0 that initiated the request. It should be understood that other GPUs (i.e., GPU1, GPU2, and GPU3) will also initiate load reduction requests for data blocks 1, 2, and 3 respectively to trigger the switch to perform corresponding operations. The specific process can refer to the above-mentioned reduction processing process for data block 0, which will not be repeated here.

[0104] S3, Broadcast: After receiving the reduction result returned by the switch, GPU0 will detect the buffer status of the local second temporary storage area. Here, GPU0 can set a Head2 memory locally to store the buffer status of the second temporary storage area, point to the memory through the HeadPtr2 pointer, and use polling to read the buffer status in the memory pointed to by the HeadPtr2 pointer. In the initial state, the buffer status of the second temporary storage area is set to buffer ready. When GPU0 reads the buffer status as buffer ready through polling, it can send a broadcast request and reduction result to the switch.

[0105] After receiving the broadcast request and reduction result, the switch broadcasts the reduction result to all GPUs (i.e., GPU0~GPU3). After each GPU receives the reduction result broadcast by the switch, it first stores it temporarily in the second temporary storage area. Subsequently, each GPU updates the buffer status of the first temporary storage area to buffer ready, and at the same time updates the data status of the second temporary storage area to data ready.

[0106] Similarly, after receiving the reduction result of data block 1 returned by the switch, GPU1 stores it in the TLR register, and when it detects that the buffer status of the second temporary storage area is buffer ready, it sends a broadcast request and the reduction result of data block 1 to the switch, so that the switch broadcasts the reduction result to all GPUs. GPU2 and GPU3 will also perform similar operations for the reduction results of data block 2 and data block 3 respectively, which will not be elaborated here.

[0107] S4, Gather: After each GPU stores the reduction result broadcast by the switch in the second temporary storage area, it detects the data status of the second temporary storage area. For example, GPU0 can set a Tail2 memory locally to store the data status of the second temporary storage area, and point to this memory through the TailPtr2 pointer. When GPU0 polls and reads that the data status is data ready, it can move the reduction result in the second temporary storage area to the output buffer in a unicast manner. Subsequently, GPU0 can update the buffer status of the second temporary storage area to buffer ready, so that the second temporary storage area can receive and store new reduction results.

[0108] It should be noted that when each GPU updates or reads the buffer status and data status of the local first temporary storage area and second temporary storage area, it can be implemented by the method of performing operations separately for each data block, or by the method of operating on the entire temporary storage area. The embodiments of the present invention do not make specific limitations on this.

[0109] It can be understood that the above four stages (Scatter, Reduce, Broadcast, Gather) are respectively executed in parallel by several warps, and these steps can be partially overlapped to achieve parallel processing. For example, when GPU0 is processing the Scatter operation of data block 2, the switch is processing the Reduce operation of data block 1, and at the same time GPU1 has already received the broadcast result of data block 0. By parallel execution and overlap of each stage, the total time of the entire computing process can be significantly shortened, thereby improving the collective communication efficiency.

[0110] Next, the collective communication device provided by the present invention will be described. The collective communication device described below can be correspondingly referred to the collective communication method described above.

[0111] Based on any of the above embodiments, Figure 6 is one of the schematic structural diagrams of the collective communication device provided by the present invention. As Figure 6 shown, this device is applied to a computing device, and the device includes: A data splitting unit 610, configured to split the data in the input buffer, and write each split data block into the first temporary storage area; A reduction request unit 620, configured to send a load reduction request to a network device, where the load reduction request is used to trigger the network device to load target data blocks from the first temporary storage area of each computing device in the communication group, and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; A broadcast request unit 630, configured to, after receiving the reduction result returned by the network device, send a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group; A result writing unit 640, configured to store the reduction result broadcast by the network device received into the second temporary storage area, and write the reduction result in the second temporary storage area to the output buffer.

[0112] The device provided by the embodiment of the present invention sends a load reduction request and a broadcast request to the network device through the computing device, and the network device uniformly processes data loading, reduction calculation, and result broadcasting. This not only reduces the computing tasks of the computing device itself, enables the computing device to focus more on the core computing logic, and improves the overall computing efficiency, but also avoids frequent and large direct data exchanges between computing devices, effectively reduces the network traffic, reduces the communication links and data transmission paths, alleviates the network bandwidth pressure, thereby reducing the communication delay and improving the collective communication efficiency. By setting the first temporary storage area and the second temporary storage area on each computing device, it can not only ensure the efficient transfer and processing of data between the computing device and the network device, but also play a certain buffering role, smooth the burst traffic of communication, and avoid network congestion. Among them, the first temporary storage area is used to temporarily store each data block written from the input buffer, so that the computing device can complete data splitting and temporary storage locally without waiting for the cooperation of the network device or other computing devices, thereby avoiding the computing device from being idle due to waiting for communication to complete; similarly, the second temporary storage area allows the computing device to temporarily store the reduction result after receiving the broadcast, and then independently write the result to the output buffer later, avoiding the output buffer from being frequently occupied or blocked. In addition, the network device will broadcast the reduction result to each computing device in the communication group only after receiving the broadcast request from the computing device. This design of actively triggering the broadcast operation by the computing device makes the broadcast operation process controllable, avoids resource conflicts, and at the same time reduces the control logic complexity of the network device.

[0113] Based on any of the above embodiments, the data splitting unit 610 is specifically configured to: When it is detected that the buffer status of the first temporary storage area is buffer ready, write each split data block into the first temporary storage area, and update the data status of the first temporary storage area to data ready; Correspondingly, the reduction request unit 620 is specifically configured to: When it is detected that the data status of the first temporary storage area is data ready, send a load reduction request to the network device.

[0114] Based on any of the above embodiments, the broadcast request unit 630 is specifically configured to: When it is detected that the buffer status of the second temporary storage area is buffer ready, send a broadcast request and the reduction result to the network device.

[0115] Based on any of the above embodiments, the result writing unit 640 is specifically configured to: After storing the reduction result broadcast by the network device received into the second temporary storage area, update the buffer status of the first temporary storage area to buffer ready, and update the data status of the second temporary storage area to data ready; When it is detected that the data status of the second temporary storage area is data ready, write the reduction result in the second temporary storage area to the output buffer, and update the buffer status of the second temporary storage area to buffer ready.

[0116] Based on any of the above embodiments, the number of loops required for all computing devices in the communication group to execute tasks is determined based on the amount of data to be processed by each thread block and the amount of data processed by each computing device in one loop. The amount of data to be processed by each thread block is determined based on the total amount of data to be processed and the number of thread blocks. The amount of data processed by each computing device in one loop is determined based on the data block size and the number of devices in the communication group.

[0117] Based on any of the above embodiments, Figure 7 is the second schematic structural diagram of the collective communication device provided by the present invention. As Figure 7 shown, this device is applied to a network device, and this device includes: A data loading unit 710, configured to receive a load reduction request sent by any computing device, and based on the load reduction request, load target data blocks from the first temporary storage area of each computing device in the communication group. The target data blocks are written into the first temporary storage area from the input buffer after each computing device splits the data in the input buffer. A reduction execution unit 720 is configured to perform a reduction operation based on each loaded target data block, obtain a reduction result, and return the reduction result to any one of the computing devices, so that any one of the computing devices sends a broadcast request and the reduction result to the network device; A result broadcasting unit 730 is configured to receive the broadcast request and the reduction result sent by any one of the computing devices, and based on the broadcast request, broadcast the reduction result to each computing device in the communication group, so that each computing device stores the reduction result in a second temporary storage area, and writes the reduction result in the second temporary storage area to an output buffer.

[0118] The device provided by the embodiment of the present invention sends a load reduction request and a broadcast request to the network device through a computing device. The network device uniformly processes data loading, reduction calculation, and result broadcasting, which not only reduces the computing tasks of the computing device itself, enables the computing device to focus more on the core computing logic, and improves the overall computing efficiency, but also avoids frequent and large direct data exchanges between computing devices, effectively reduces the network traffic, reduces the communication links and data transmission paths, alleviates the network bandwidth pressure, thereby reducing the communication delay and improving the collective communication efficiency. By setting a first temporary storage area and a second temporary storage area on each computing device, it can not only ensure the efficient transfer and processing of data between the computing device and the network device, but also play a certain buffering role, smooth the burst traffic of communication, and avoid network congestion. Among them, the first temporary storage area is used to temporarily store each data block written from the input buffer, so that the computing device can complete data segmentation and temporary storage locally without waiting for the cooperation of the network device or other computing devices, thereby avoiding the idle state of the computing device due to waiting for communication to complete; similarly, the second temporary storage area allows the computing device to temporarily store the reduction result after receiving the broadcast reduction result, and then independently write the result to the output buffer later, avoiding the frequent occupation or blockage of the output buffer. In addition, the network device will broadcast the reduction result to each computing device in the communication group only after receiving the broadcast request from the computing device. This design of actively triggering the broadcast operation by the computing device makes the broadcast operation process controllable, avoids resource conflicts, and at the same time reduces the control logic complexity of the network device.

[0119] Figure 8 is a schematic structural diagram of the device provided by the present invention, as Figure 8As shown, the device can be a computing device or a network device. The device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute a collective communication method, which is applied to a computing device. The method includes: splitting the data in the input buffer and writing each split data block into the first temporary storage area; sending a load reduction request to the network device, where the load reduction request is used to trigger the network device to load target data blocks from the first temporary storage area of each computing device in the communication group, and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; after receiving the reduction result returned by the network device, sending a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group; storing the reduction result broadcast by the network device received into the second temporary storage area, and writing the reduction result in the second temporary storage area to the output buffer.

[0120] In addition, the processor 810 can also call the logical instructions in the memory 830 to execute a collective communication method, which is applied to a network device. The method includes: receiving a load reduction request sent by any computing device, and based on the load reduction request, loading target data blocks from the first temporary storage area of each computing device in the communication group, where the target data blocks are split from the data in the input buffer by each computing device and written into the first temporary storage area from the input buffer; performing a reduction operation based on the loaded target data blocks to obtain a reduction result, and returning the reduction result to the any computing device, so that the any computing device sends a broadcast request and the reduction result to the network device; receiving the broadcast request and the reduction result sent by the any computing device, and based on the broadcast request, broadcasting the reduction result to each computing device in the communication group, so that each computing device stores the reduction result in the second temporary storage area, and writes the reduction result in the second temporary storage area to the output buffer.

[0121] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0122] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the collective communication method provided by the above-mentioned various methods. This method is applied to a computing device and includes: splitting the data in the input buffer and writing each split data block into a first temporary storage area; sending a load reduction request to a network device, where the load reduction request is used to trigger the network device to load target data blocks from the first temporary storage area of each computing device in the communication group and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; after receiving the reduction result returned by the network device, sending a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group; storing the reduction result broadcast by the network device received into a second temporary storage area and writing the reduction result in the second temporary storage area to the output buffer.

[0123] In addition, when the computer program is executed by a processor, the computer is further capable of executing the collective communication method provided by each of the above methods. This method is applied to a network device and includes: receiving a load reduction request sent by any computing device, and based on the load reduction request, loading target data blocks from the first scratchpad area of each computing device in the communication group. The target data blocks are obtained by each computing device splitting the data in the input buffer and then writing the data from the input buffer to the first scratchpad area; performing a reduction operation based on the loaded target data blocks to obtain a reduction result, and returning the reduction result to the any computing device, so that the any computing device sends a broadcast request and the reduction result to the network device; receiving the broadcast request and the reduction result sent by the any computing device, and based on the broadcast request, broadcasting the reduction result to each computing device in the communication group, so that each computing device stores the reduction result in the second scratchpad area and writes the reduction result in the second scratchpad area to the output buffer.

[0124] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the collective communication method provided by each of the above methods. This method is applied to a computing device and includes: splitting the data in the input buffer and writing each split data block to the first scratchpad area; sending a load reduction request to the network device, where the load reduction request is used to trigger the network device to load target data blocks from the first scratchpad area of each computing device in the communication group and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; after receiving the reduction result returned by the network device, sending a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group; storing the received reduction result broadcast by the network device in the second scratchpad area and writing the reduction result in the second scratchpad area to the output buffer.

[0125] In addition, when the computer program is executed by a processor, it implements a collective communication method provided by the above-mentioned various methods. This method is applied to a network device and includes: receiving a load reduction request sent by any computing device, and based on the load reduction request, loading target data blocks from the first scratchpad area of each computing device in the communication group. The target data blocks are written from the input buffer to the first scratchpad area after each computing device slices the data in the input buffer; based on the loaded target data blocks, performing a reduction operation to obtain a reduction result, and returning the reduction result to the any computing device, so that the any computing device sends a broadcast request and the reduction result to the network device; receiving the broadcast request and the reduction result sent by the any computing device, and based on the broadcast request, broadcasting the reduction result to each computing device in the communication group, so that each computing device stores the reduction result in the second scratchpad area and writes the reduction result in the second scratchpad area to the output buffer.

[0126] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A collective communication method, characterized in that, The method is applied to a computing device, and the method includes: Segment the data in the input buffer and write each segmented data block into the first temporary storage area; Send a load reduction request to the network device, where the load reduction request is used to trigger the network device to load target data blocks from the first temporary storage area of each computing device in the communication group, and perform a reduction operation based on the loaded target data blocks to obtain a reduction result; After receiving the reduction result returned by the network device, send a broadcast request and the reduction result to the network device, where the broadcast request is used to trigger the network device to broadcast the reduction result to each computing device in the communication group; Store the reduction result broadcast by the network device received into the second temporary storage area, and write the reduction result in the second temporary storage area to the output buffer.

2. The collective communication method according to claim 1, wherein The writing each segmented data block into the first temporary storage area includes: When it is detected that the buffer status of the first temporary storage area is buffer ready, write each segmented data block into the first temporary storage area, and update the data status of the first temporary storage area to data ready; Correspondingly, the sending a load reduction request to the network device includes: When it is detected that the data status of the first temporary storage area is data ready, send a load reduction request to the network device.

3. The set communication method according to claim 1, wherein The sending the broadcast request and the reduction result to the network device includes: When it is detected that the buffer status of the second temporary storage area is buffer ready, send a broadcast request and the reduction result to the network device.

4. The collective communication method according to claim 1, characterized in that, After storing the reduction result broadcast by the network device received into the second temporary storage area, it further includes: Update the buffer status of the first temporary storage area to buffer ready, and update the data status of the second temporary storage area to data ready; Correspondingly, the writing the reduction result in the second temporary storage area to the output buffer includes: When it is detected that the data status of the second temporary storage area is data ready, write the reduction result in the second temporary storage area to the output buffer, and update the buffer status of the second temporary storage area to buffer ready.

5. The set communication method according to any one of claims 1 to 4, characterized in that The number of loop times required for all computing devices in the communication group to execute the task is determined based on the data volume to be processed by each thread block and the data volume processed by each computing device in one loop. The data volume to be processed by each thread block is determined based on the total data volume to be processed and the number of thread blocks. The data volume processed by each computing device in one loop is determined based on the data block size and the number of devices in the communication group.

6. A collective communication method, characterized in that, The method is applied to a network device, and the method includes: Receive a load reduction request sent by any computing device. Based on the load reduction request, load target data blocks from the first temporary storage area of each computing device in the communication group. The target data blocks are obtained by each computing device segmenting the data in the input buffer and writing the data from the input buffer into the first temporary storage area; Based on each loaded target data block, perform a reduction operation to obtain a reduction result, and return the reduction result to any one of the computing devices, so that any one of the computing devices sends a broadcast request and the reduction result to the network device; Receive the broadcast request and the reduction result sent by any one of the computing devices, and based on the broadcast request, broadcast the reduction result to each computing device in the communication group, so that each computing device stores the reduction result in the second temporary storage area and writes the reduction result in the second temporary storage area to the output buffer.

7. A computing device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the collective communication method according to any one of claims 1 to 5.

8. A network device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the computer program, it implements the collective communication method according to claim 6.

9. A collective communication system, characterized in that, It includes a plurality of computing devices according to claim 7, and a network device according to claim 8.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the collective communication method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Data processing method and device, electronic equipment and storage medium

    CN121957913A

  • Communication method and device, electronic equipment, computer readable storage medium and computer program product

    CN122019456A

  • Communication method and apparatus, electronic device, computer-readable storage medium, and computer program product

    CN122019456B