System and method for reducing delay of summing of multiple GPU card protocols in distributed text reasoning

By adopting the data packaging and reception confirmation mechanism in the process of multi-GPU card specification summing, the atomic visibility and sequential visibility of the CUDA memory model is used to solve the problem of high latency for multi-GPU card specification summing, and more efficient synchronization and data transmission are achieved.

CN120407180AActive Publication Date: 2025-08-01BEIJING SILICONFLOW TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510515280.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-01
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The prior art has high latency problems in the process of summing the multi-GPU card specifications, especially in scenarios with small communication volume, the synchronization operation takes a long time, resulting in low overall efficiency.

Method used

The data packaging component is used to package the data according to a predetermined size, including shard data and flag bit data, and is transmitted to the buffer of the remote GPU card through the data transmission component, and data reception and confirmation is used to use the atomic visibility and sequential visibility of the CUDA memory model to ensure that all GPU cards are synchronized.

Benefits of technology

It significantly reduces the delay of multi-GPU card specification summing, improves bandwidth utilization, reduces the overhead of synchronous operation, makes full use of the parallel characteristics of GPU, and improves the efficiency of multi-card specification summing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407180A_ABST
    Figure CN120407180A_ABST
Patent Text Reader

Abstract

The invention relates to a system for reducing delay of summing of multiple GPU card protocols in distributed text reasoning and a method thereof. The system comprises: a data packaging component, which packages the data of a GPU card, which needs to execute a multi-GPU card protocol locally, according to a predetermined size, a first part of each data packet being data to be transmitted segmented in sequence from fragmented data, and a second part being flag bit data; the data transmission component is used for transmitting the packaged data packet from the local GPU card to buffers of other GPU cards at the far end; and the data receiving confirmation component is used for completing polling aiming at corresponding identification bit data when receiving of a data packet with a preset size is completed in the polling waiting execution process, and enabling a stage receiving completion result to be seen by all GPU cards based on the visibility of data transmission and to be confirmed, so that multi-card protocol summation is performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer information processing, and more particularly, to a system and method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning. Background Art

[0002] With the explosion of ChatGPT, large models have quickly attracted the attention of the contemporary Internet industry, and the demand for the use of large language models in various industries has also increased accordingly. The demand for text generation has also increased exponentially. However, text generation itself consumes a large amount of computing power. How to deploy and accelerate the inference of large language models as much as possible in the currently widely used device environment is a hot research topic.

[0003] Model parallelism, a commonly used parallelism method in large model training, can naturally be applied to the text generation inference of large prediction models. Among them, AllReduce (global reduction or all-reduction) multi-card reduction summation is its main communication method. Currently, the mainstream deep learning frameworks in the industry all follow the NVIDIA communication library NCCL and call its AllReduce implementation. However, in some scenarios with small communication volume and latency sensitivity, NCCL has no targeted optimization, and thus the communication time accounts for a high proportion.

[0004] The two inference frameworks, TensorRTT (Tensor RunTime) and vLLM (Virtual Large Language Model), have each implemented Custom AllReduce for scenarios with small communication volume. When GPU supports P2P direct access, through the CUDA IPC inter-process communication mechanism, each GPU can obtain the data pointer on other GPUs, and then multi-card reduction summation can be achieved. During this process, it is also necessary to insert GPUBarrier (GPUBarrier, GPU barrier, a synchronization mechanism used to coordinate the execution order of multiple threads / processing units in parallel computing. By forcing waiting for all associated operations to complete before continuing to execute, it ensures data consistency and the correctness of the execution logic) to achieve synchronization between GPUs.

[0005] In multi-GPU card or distributed training scenarios, the global Barrier is used to synchronize the computing progress of different devices. For example, torch.distributed.barrier() in PyTorch ensures that all processes start training only after completing data loading. In scenarios with low communication volume, the actual reduction summation operation has a relatively low computational intensity and takes less time. However, there is a certain overhead in GPU synchronization, and the GPU synchronization time may be comparable to the time taken by the actual reduction summation operation. Therefore, the key to achieving efficient low-latency multi-GPU reduction summation lies in an efficient GPU Barrier. In each reduction summation, a GPU Barrier needs to be inserted at the beginning and end. The first GPU Barrier ensures that all GPUs have executed to the AllReduce Kernel, while the second GPU Barrier ensures that all GPUs have completed reading the data on other GPUs. Otherwise, the subsequent operations may corrupt the data on the GPUs, resulting in incorrect multi-GPU reduction summation results.

[0006] In the implementation of vLLM, its GPU Barrier is implemented by writing a Flag bit to the remote GPU locally, and locally polling (busy wait) to read the bit until it detects that all remote GPUs have completed writing the bit, indicating that all GPUs have completed synchronization. Although, with P2P support, the overhead of writing the bit to the remote GPU is very low (~1us), as the number of thread blocks launched by the GPU increases, the overall GPU synchronization overhead also increases. Therefore, in the implementation of vLLM, there is a certain limit on the number of thread blocks launched by the GPU. One drawback of limiting the number of thread blocks is that when the number of reduction summations is large, the load of each thread block (In the same thread block, the Barrier ensures that all threads pause after executing to a specified position until all threads reach that point and then continue with subsequent operations. For example, the syncthreads() instruction in CUDA is used to synchronize the reading and writing of shared memory) is high, and it is not possible to fully utilize the parallel characteristics of the GPU, making the entire reduction summation process in a high-latency state and not very efficient.

[0007] Therefore, it is expected to at least reduce the latency length of multi-GPU card reduction summation in existing distributed text inference, thereby improving the efficiency of multi-GPU reduction summation.

[0008] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0009] In view of this, the present disclosure provides a distributed text inference system for large language models, which can at least reduce the current high latency state of multi-card reduction summation.

[0010] Other features and advantages of the present disclosure will become apparent from the following detailed description, or will be learned in part through the practice of the present disclosure.

[0011] According to an aspect of the present disclosure, a system for reducing the latency of multi-GPU card reduction summation in distributed text inference is proposed, including: a data packing component that packs the data that needs to perform multi-GPU card reduction on the local GPU card according to a predetermined size, where the first part of each data packet is the data to be transmitted sequentially segmented from the sharded data, and the second part is the flag bit data; a data transmission component that transmits the packed data packets from the local GPU card to the buffer of other remote GPU cards; and a data reception confirmation component that, during the polling waiting process, when the reception of a data packet of a predetermined size is completed, polls for the completion of the corresponding flag bit data, and based on the visibility of data transmission, makes the staged reception completion result visible and confirmed to all GPU cards for multi-card reduction summation.

[0012] According to the system for reducing the latency of multi-GPU card reduction summation in distributed text inference of the present disclosure, the first part of each data packet is 32-bit data to be transmitted sequentially segmented from the sharded data, the second part is 32-bit flag bit data, and the data transmission component transmits double data packets from the local GPU card to the buffer of other remote GPU cards each time, and the data reception confirmation component of each GPU card polls for the completion of the corresponding flag bit data when the reception of a 64-bit data packet is completed during the polling waiting process, and based on the 64-bit atomic visibility in the CUDA memory model, makes the staged reception completion result visible and confirmed to all GPU cards for multi-card reduction summation.

[0013] According to the system for reducing the latency of multi-GPU card reduction summation in distributed text inference of the present disclosure, the first part of each data packet is 120-byte data to be transmitted sequentially segmented from the sharded data, the second part is 8-byte flag bit data, and the data transmission component transmits the data packets from the local GPU card to the buffer of other remote GPU cards based on the nvlink data transmission channel each time, and the data reception confirmation component of each GPU card performs a polling waiting process to poll for the completion of the corresponding flag bit data when the reception of a 128-byte data packet is completed, and based on the sequential visibility in the CUDA memory model, makes the staged reception completion result visible and confirmed to all GPU cards for multi-card reduction summation.

[0014] A system for reducing the latency of multi-GPU card reduction summation in distributed text inference according to the present disclosure, wherein the data transmission component of each GPU card respectively transmits continuous packets of the local full amount of data from the local GPU card to the buffers of other remote GPU cards at one time, so that each GPU card performs multi-card reduction summation locally.

[0015] A system for reducing the latency of multi-GPU card reduction summation in distributed text inference according to the present disclosure, wherein the data packing component of each GPU card evenly divides the sharded data into sub-sharded data according to the number of GPU cards before packing, and then packs each sub-sharded data, and the data transmission component only transmits the continuous packets of the corresponding sub-sharded data from the local GPU card to the buffers of the corresponding remote GPU cards respectively, and after each GPU card receives all its corresponding sub-sharded data and performs multi-card reduction summation on all the received sub-sharded data locally to obtain the reduced summation result sub-sharded data, each sends its local reduced summation result sub-sharded data to other GPU cards.

[0016] According to another aspect of the present disclosure, a method for reducing the latency of multi-GPU card reduction summation in distributed text inference is provided, including: packing the data that needs to perform multi-GPU card reduction on the local GPU card by a data packing component according to a predetermined size, wherein the first part of each data packet is the data to be transmitted sequentially segmented from the sharded data, and the second part is the flag bit data; transmitting the packed data packets from the local GPU card to the buffers of other remote GPU cards by a data transmission component; and performing a polling waiting process by a data reception confirmation component, so that when the reception of the data packets of the predetermined size is completed, the polling for the corresponding flag bit data is completed, and based on the visibility of the data transmission, the staged reception completion result is visible and confirmed to all GPU cards, so as to perform multi-card reduction summation.

[0017] A method for reducing the latency of multi-GPU card reduction summation in distributed text inference according to the present disclosure, wherein the first part of each data packet is 32-bit data to be transmitted sequentially segmented from the sharded data, the second part is 32-bit flag bit data, and the data transmission component transmits double data packets from the local GPU card to the buffers of other remote GPU cards each time, and the data reception confirmation component of each GPU card performs a polling waiting process, so that when the reception of the 64-bit data packets is completed, the polling for the corresponding flag bit data is completed, and based on the 64-bit atomic visibility in the CUDA memory model, the staged reception completion result is visible and confirmed to all GPU cards, so as to perform multi-card reduction summation.

[0018] Method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning according to the present disclosure, wherein the first part of each data packet is 120 bytes of data to be transmitted sequentially segmented from the sharded data, the second part is 8 bytes of flag bit data, and the data transmission component transfers the data packet from the local GPU card to the buffer of other remote GPU cards each time based on the nvlink data transmission channel, and the data reception confirmation component of each GPU card performs a polling wait process, so that when the reception of the 128-byte data packet is completed, the polling for the corresponding identification bit data is completed, and the staged reception completion result is made visible and confirmed to all GPU cards based on the sequential visibility in the CUDA memory model, so as to perform multi-card reduction summation.

[0019] Method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning according to the present disclosure, wherein the data transmission component of each GPU card transfers the continuously packed data packets of the local full amount of data from the local GPU card to the buffers of other remote GPU cards respectively, so that each GPU card performs multi-card reduction summation locally.

[0020] Method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning according to the present disclosure, wherein the data packing component of each GPU card evenly divides the sharded data into sub-sharded data corresponding to the number of GPU cards before packing, and then packs each sub-sharded data, and the data transmission component only transfers the continuous data packets of the corresponding sub-sharded data from the local GPU card to the buffers of the corresponding remote GPU cards respectively, and after each GPU card receives all its corresponding sub-sharded data and performs multi-card reduction summation on all the received sub-sharded data locally to obtain the reduced summation result sub-sharded data, each sends its local reduced summation result sub-sharded data to other GPU cards.

[0021] Systems and methods for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning according to the present disclosure. Considering that in the CUDA memory model, 64-bit atomic visibility is available, by combining 32-bit data and a 32-bit Flag bit into a 64-bit data packet each time, when the 32-bit Flag bit can be verified, it also means that the previous 32-bit data has been transmitted. Each thread in CUDA supports a maximum of 128-bit read and write, so two data packets can be written into the remote buffer at one time. The remote GPU reads the Buffer through busy wait and verifies whether it is the value of the Flag bit. When it is equal to the bit value, it means that the data has been successfully transmitted. The present disclosure utilizes the characteristics that CUDA memory operations are visible in sequence for 128 Bytes and each thread in CUDA supports a maximum of 16 Bytes read and write. Therefore, the data to be transmitted is grouped into 128 Bytes as a data group, and the data block is divided into 8 threads with 16 Bytes for each thread, so that each group just meets the 128-Byte transmission limit of CUDA. For the last thread of each thread group, it stores 8 Bytes of data and an 8-Byte Flag bit. The 8-Byte Flag bit is exactly 64 bits, meeting the characteristic of 64-bit atomic visibility in the CUDA memory model. Thus, the data packet verification is only completed by this thread, and the other threads in the thread group store 16 Bytes of valid data. Therefore, when the Flag bit in the last thread of each thread group meets the expected value, it means that the previous 120 Bytes of data have been successfully transmitted. This greatly utilizes the low-latency and high-bandwidth optimization protocol designed for collective communication between GPUs (such as All Reduce, Broadcast) in NVIDIA NCCL (Cluster Communication Library), making the protocol bandwidth utilization rate reach more than 90%, and also improving the rate of multi-card reduction summation and reducing its latency from the perspective of bandwidth utilization. In addition, by reducing the data transmission volume during the reduction summation through two-stage reduction summation, the latency of the reduction summation is also reduced.

[0022] It should be understood that the above general description and the following detailed description are exemplary only and do not limit the present disclosure. Brief Description of the Drawings

[0023] By referring to the accompanying drawings and describing its exemplary embodiments in detail, the above and other objects, features, and advantages of the present disclosure will become more apparent. The following described drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a block diagram of the first embodiment of a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning according to an exemplary embodiment.

[0025] Figure 2 Shown is a schematic diagram of the process of multi-card reduction summation by a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning.

[0026] Figure 3 Shown is a comparison diagram of the test effects of a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning according to the present disclosure.

[0027] Figure 4 Shown is a schematic diagram of the second implementation manner of the process of multi-card reduction summation by a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning.

[0028] Figure 5 Shown is a comparison diagram of the test effects of a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning based on the LL128 protocol according to the present disclosure.

[0029] Figure 6 Shows a schematic diagram of the process of one-stage multi-card reduction summation by a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning according to the present disclosure.

[0030] Figure 7 Shown is a schematic diagram of the process of two-stage multi-card reduction summation by a system for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning according to the present disclosure.

[0031] Figure 8 Shown is a comparison diagram of the test effects of the first-order and second-order systems for reducing the latency of reduction summation of multiple GPU cards in distributed text reasoning according to the present disclosure. Detailed implementation manners

[0032] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.

[0033] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0034] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0035] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0036] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first computing device discussed below may be referred to as the second computing device without departing from the teachings of the concepts of the present disclosure. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0037] Those skilled in the art can understand that the drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present disclosure, so they cannot be used to limit the protection scope of the present disclosure.

[0038] Figure 1 is a block diagram of a first embodiment of a system 100 for reducing the latency of reduction summation of multiple GPU cards in distributed text inference shown according to an exemplary embodiment. As Figure 1As shown, the system 100 for reducing the latency of multi-GPU card reduction summation in distributed text inference includes: a data packing component 110 that packs the data that needs to perform multi-GPU card reduction locally on the GPU card according to a predetermined size, where the first part of each data packet is the data to be transmitted sequentially segmented from the shard data, and the second part is the flag bit data; a data transmission component 120 that transmits the packed data packets from the local GPU card to the buffers of other remote GPU cards; and a data reception confirmation component 130 that, during the polling wait process, when the reception of a data packet of a predetermined size is completed, polls for the completion of the corresponding identification bit data, and makes the staged reception completion result visible and confirmed to all GPU cards based on the visibility of the data transmission, so as to perform multi-card reduction summation. Although Figure 1 the data packing components 110 in

[0039] are respectively labeled as 110-1, 110-2... 110-N, when not distinguishing from each other, they are uniformly referred to as 110. Similarly, when the data transmission component 120 and the data reception confirmation component 130 are in different GPU cards, they are also referred to by the simplified labels 120 and 130. Figure 1 As shown, the data packing component 110 for each GPU card in the system 100 for reducing the latency of multi-GPU card reduction summation in the present disclosure utilizes 64-bit atomic visibility in the CUDA memory model, and combines 32-bit size data and 32-bit size Flag flag bits into a 64-bit data packet each time. When the GPU card that receives the data packet, the data reception confirmation component 130 reads the data packet in the buffer through busy wait to verify the 32-bit Flag flag bit. When the verification of the 32-bit Flag flag bit passes, it also means that the previous 32-bit data has also been transmitted. Figure 2The figure shows a schematic diagram of the multi-card reduction summation process of the system for reducing the latency of multi-GPU card reduction summation in distributed text inference. It should be noted that each thread in CUDA supports a maximum of 128-bit read and write. Therefore, the data transmission component 120 of the present disclosure can write two data packets into the buffer of the remote GPU card at one time. The data reception confirmation component 130 of the remote GPU card reads the data buffer through busy wait. For example, the data reception confirmation component 130-2 of GPU2 polls whether the newly received data in the data buffer 2 contains a 32-bit Flag flag bit and checks whether it is the value of the Flag flag bit. When it is equal to the flag bit value, it indicates that the valid data corresponding to the flag bit has been successfully transmitted. The 64-bit atomic visibility is an atomic operation for 64-bit data types that not only ensures the atomicity of the operation but also ensures through hardware mechanisms that the operation result is immediately visible to other threads. The atomic operation for 64-bit data types means that in a multi-threaded environment, a single instruction uninterruptedly completes the read-write-modify-write-back process of memory, avoiding data race, and "visible" means the visibility of the operation result to other threads, which is an extended feature of atomic operations. When a certain thread completes an atomic operation, other threads must immediately observe the latest value when accessing the same memory location without explicit synchronization instructions. Figure 2 The figure shows a schematic diagram of the first embodiment of the multi-card reduction summation process of the system for reducing the latency of multi-GPU card reduction summation in distributed text inference. As Figure 2 shown, each GPU, such as GPU0, continuously packs the local shard data to be reduced and summed into the first part of valid data D and the second part of flag data F, thus forming continuous data packets D0-1+F0-1, D0-2+F0-2, and so on. Subsequently, these are directly written into the buffers of other GPUs, such as buffer 0, buffer 1, buffer 2, and buffer 3, through memory access. The data reception confirmation components 130 of the remote GPUs GPU1, GPU2, and GPU3 perform busy wait and make the polling verification results atomically visible so that other GPUs can know, achieving synchronization. This eliminates the defect in the prior art where a barrier ensures that all threads pause after executing to a specified position and at the same time ensures the integrity of all data.

[0040] As described above, the present invention mainly uses this technology to accelerate the reduction summation process in order to prevent the delay caused by the mutual waiting of the existing multi-card protocols for summation, but uses this technology to solve the delay problem of multi-card reduction summation. This 64-bit atomic visibility itself belongs to the prior art and will not be elaborated in detail here. Since the present disclosure combines data of 32-bit size and a 32-bit Flag bit into a 64-bit data packet each time by using 64-bit atomic visibility, the overhead of the entire GPU synchronization is significantly reduced, and the limit on the number of thread blocks started by the GPU is basically eliminated.

[0041] Figure 3 The figure shows a comparison of the system test effects of reducing the reduction summation delay of multiple GPU cards in distributed text inference according to the present disclosure. As Figure 3 shown, the distributed data processing system uses 8 H800 GPU cards, adopts different reduction summation systems, and tests the reduction summation time for different amounts of transmitted data that need to be reduced and summed. As Figure 3 shown, for the same amount of transmitted data, the existing NCCL reduction summation takes the most time, that is, the longest delay time, the vLLM customized reduction summation method takes the second longest time, and the cross-node customized reduction summation method adopted in the present application takes the shortest time, that is, the smallest reduction summation delay time.

[0042] Although the above-mentioned packing method can be used to reduce the reduction summation delay of multiple cards under the LL protocol, however, under the GPU synchronization mechanism based on the LL protocol, half of each transmitted data packet is used to store the Flag bit, and the other half stores the valid data, which means that half of the bandwidth is wasted during the transmission process. As the communication volume increases, the multi-card reduction summation operation gradually changes from latency-bound to memory-bound, and the effective bandwidth utilization rate of the LL protocol is not high, that is, the utilization efficiency of the bandwidth is only half. Objectively, this means that in this way, if the existing bandwidth is fully utilized, the delay can be further reduced. The multi-card reduction summation synchronization mechanism based on the LL128 protocol for large amounts of data is even more inefficient. The LL128 protocol is a low-latency and high-bandwidth optimized protocol designed for inter-GPU collective communication (such as AllReduce, Broadcast) in NVIDIA NCCL (cluster communication library). Of course, in the case of reduction summation of small amounts of data, under the LL protocol mechanism, Figure 2 the shown method can significantly reduce the delay effect.

[0043] To make full use of the high-bandwidth optimized protocol based on LL128, Figure 4Schematic diagram of the second implementation manner of the multi-card reduction summation process in the system for reducing the latency of multi-GPU card reduction summation in distributed text reasoning. To this end, the Figure 4 shown method is adopted to fully utilize the bandwidth utilization. As Figure 4 shown, when nvlink is configured inside the GPU cluster adopted by the deployed distributed data processing system in the present disclosure, its memory operation is sequentially visible in 128 bytes (Bytes). Nvlink is a high-speed interconnection technology developed by NVIDIA, dedicated to improving the communication efficiency between the GPU and the CPU, between the GPU and other devices (such as memory, other GPUs), especially serving high-performance computing (HPC) and artificial intelligence (AI) scenarios. Here, only this technology is utilized, rather than improving this technology, so it will not be elaborated. The present disclosure utilizes the sequential visibility of 128 Bytes in memory operation when nvlink is configured inside the GPU cluster, and splits the 128-Byte data packet into 120-Byte data and 8-Byte flag bits. Since each thread in CUDA supports a maximum of 16 Bytes of reading and writing, the sharded data in the present disclosure is packed in accordance with 128 Bytes, and the 128 Bytes are divided into 8 thread groups according to the maximum 16 Bytes supported by each thread, so that each group just meets the 128-Byte transmission limit. For the last thread of each thread group, the present disclosure makes it store 8 bytes (Bytes) of data and 8-byte (Bytes) Flag flag bits, and the data packet verification is only completed by the last thread, while the other threads in the thread group all store 16 Bytes of valid data. Therefore, as Figure 4As shown in the figure, the data packaging component 110 packs the sharded data to be reduced and summed according to the above 128 bytes, and groups them into 8 groups according to the maximum supported size of 16 bytes per thread within the packet. In this way, the first part of each data packet is 120 bytes of data to be transmitted sequentially segmented from the sharded data. These 120 bytes of valid data are divided into 7.5 threads, including data D0-0 - D0-14. The valid data D14 of the last half thread is combined with the 8-byte flag data F0- in the second part to form a thread 8. The data transmission component 120 each time transfers the data packets from the local GPU0 card to the buffer 1-3 of other remote GPU cards in the order of thread groups based on the nvlink data transmission channel, and the data reception confirmation component 130 of each GPU card performs a polling wait process. So that when the reception of a 128-byte data packet is completed, polling is completed for the corresponding identification bit data F0-, and based on the sequential visibility in the CUDA memory model, the phased reception completion result is visible and confirmed by all GPU cards, so as to perform multi-card reduction and summation. The same operation is also performed on other GPU cards. When the Flag flag bit in the last thread of each thread group meets the expected value, it means that the previous 120 bytes of data have been successfully transmitted.

[0044] Adopt Figure 4 the method shown in the figure for reduction and summation, and the effective bandwidth utilization rate of the LL128 protocol is 120B / 128B = 93.75% Therefore, the bandwidth utilization rate of the LL128 protocol has been greatly improved compared with the LL protocol. Therefore, under large communication volumes, the multi-card reduction and summation time based on the LL128 protocol is lower. Figure 5 Shown in the figure is a comparison chart of the system test results for reducing the latency of multi-GPU card reduction and summation in distributed text inference based on the LL128 protocol according to the present disclosure. As Figure 5 shown in the figure, the distributed data processing system uses 8 H800 GPU cards, and adopts the LL protocol and the LL128 protocol reduction and summation protocols to test the reduction and summation time for different amounts of transmitted data that need to be reduced and summed. As Figure 5 shown in the figure, under the same amount of transmitted data, the time spent on reduction and summation based on the LL protocol is significantly higher than that based on the LL128 protocol reduction and summation protocol. This is a significant effect caused by taking advantage of the sequential visibility of 128 bytes (Bytes) of memory operations when using the nvlink configured inside the GPU cluster and fully adopting the sub-packet verification method of the bandwidth to improve the bandwidth utilization efficiency. It should be noted that here, one-shot data transmission is used for both, rather than phased reduction and summation data transmission.

[0045] The one-stage data transmission is a one-stop process. That is, when the GPU supports P2P, the latency of directly accessing the data pointers of each card to send data is relatively low. Based on this premise, the data transmission component 120 of each GPU card in the present disclosure transmits the continuously packed data packets of the local full amount of data from the local GPU card to the buffers of other remote GPU cards at one time, so that each GPU card performs multi-card reduction summation locally. This is the method of one-stage multi-card reduction summation: each GPU sends data packets to the Buffer on other GPUs based on the LL / LL128 communication protocol. When the GPU obtains the data on all ranks, it completes the reduction summation operation locally and writes it back to the result. Figure 6 shows a schematic diagram of the one-stage multi-card reduction summation process of the system for reducing the latency of multi-GPU card reduction summation in distributed text inference according to the present disclosure. As Figure 6 shown, assuming there are a total of N GPUs ( Figure 6 corresponding to 4, it can be more or less), and the amount of data on each GPU is M, then the communication volume of each GPU is: CommSize = (N - 1) * M where the local access of the GPU to the cache of the GPU does not require communication volume. Although the one-stage multi-card reduction summation can reduce the latency, since each GPU needs to transmit the local full amount of data to the GPU, the communication volume is relatively large. Therefore, the one-stage multi-card reduction summation method is only applicable to scenarios with relatively low communication volume.

[0046] Therefore, the present disclosure reduces the communication volume of each GPU through two-stage multi-card reduction summation, thereby reducing the latency of reduction summation. Figure 7 shown is a schematic diagram of the two-stage multi-card reduction summation process of the system for reducing the latency of multi-GPU card reduction summation in distributed text inference according to the present disclosure. Specifically, before packing, the data packing component 110 of each GPU card evenly divides the sharded data into sub-sharded data corresponding to the number of GPU cards for each local full amount of data, and then packs the sub-sharded data. Specifically, for example Figure 7As shown, when 4 GPU cards are distributed and deployed, the amount of data on each GPU is M. We evenly divide the data to be reduced and summed into 4 sub - slices, such as [M0, M1, M2, M3]. After evenly splitting the data for reduction and summation according to the number of GPUs, each GPU data transfer component 120 transfers consecutive data packets of the corresponding sub - slice data from the local GPU card to the buffers of the corresponding remote GPU cards respectively. For example, data transfer component 120 - 0 sends its [M1, M2, M3] to buffer 1 of GPU1 card, buffer 2 of GPU2 card, and buffer 3 of GPU3 card respectively. Data transfer component 120 - 1 sends its [M0, M2, M3] to buffer 0 of GPU0 card, buffer 2 of GPU2 card, and buffer 3 of GPU3 card respectively, and so on. During this first - stage reduction - sum data transfer process, the data transfer volume of each GPU card is: CommSize1 = (N - 1)*M / N.

[0047] In Figure 7 the case where there are only 4 GPU cards, the data transfer volume is 3M / 4. Each GUP card data transfer component 120 sends one of the sub - slices to the buffers of other corresponding GPU cards. Each GPU is only responsible for reducing and summing the data within one sub - slice. For example, GPU0 is responsible for reducing and summing the M0 data on each GPU (for 4 cards, there are 4 different M0 sub - slice data), GPU1 is responsible for reducing and summing the sub - slice M1 data on each GPU (for 4 cards, there are 4 different M1 sub - slice data), and so on. Therefore, each GPU needs to send the local data sub - slice to the corresponding GPU: GPU0 sends M0 to GPU0, M1 to GPU1, M2 to GPU2, and M3 to GPU3. At this time, the data reception confirmation component 130 of each GPU card will also perform a polling wait process, and after verifying the flag - bit data, it will make other GUP cards know the reception result through atomic visibility or sequential visibility.

[0048] After completing the reduction and summation within the sub - slice, the data transfer component 120 sends the local reduction - sum result sub - slice data to other GPU cards, so that each GPU has the complete reduction - sum result. This result transfer process is executed by the data transfer component 120, and its data transfer volume is also: CommSize2 = (N - 1)*M / N.

[0049] Therefore, the total data transfer volume in the two stages is CommSize = 2(N - 1)*M / N.

[0050] Compared with the communication volume (N - 1)*M per card in the one - stage reduction - sum method, Where N >= 2. It is not difficult to see that the communication volume of the two-stage reduction summation method is always less than or equal to that of the one-stage reduction summation method. In scenarios with a large communication volume, its communication time consumption is lower than that of the one-stage reduction summation method. It should be noted that since the two-stage reduction summation method involves two communication operations, this part of the delay will be slightly higher. In the case of an extremely small communication volume, the time consumption caused by these two communication operations will offset the advantage brought by reducing the communication volume, and sometimes it may be slightly higher than that of the one-stage reduction summation method. Figure 8 Shown is a comparison chart of the first-order and second-order system test effects of reducing the reduction summation latency of multiple GPU cards in distributed text inference according to the present disclosure. As Figure 8 shown, in the case of a relatively low communication volume, the communication latency time length of the first-order reduction summation is actually lower than that of the second-order reduction summation. Therefore, in a distributed processing system, although theoretically the communication volume of the second-order reduction summation is lower than that of the first-order reduction summation, its actual reduction summation latency is negatively affected due to the existence of two communication operations in the second-order reduction summation. However, in the case of a relatively large communication data volume, the effect brought by the reduction of the communication volume in the second-order reduction summation will be higher than the negative effect of the two communication operations, thus significantly reducing the overall latency time.

[0051] Optionally, since multi-card reduction summation requires copying the input data into a buffer created in advance using the cudaIPC inter-process communication mechanism, the data and flag bits will be written into the buffer. In order to distinguish each multi-card reduction summation, after the reduction summation is completed, the buffer needs to be reset or the flag bit incremented. Without these processes, when entering the next reduction summation operation, because the flag bit of the previous reduction summation is written into the buffer, the GPU will mistakenly think that the synchronization between multiple cards has been completed, resulting in an incorrect reduction summation result. In response to this situation, the existing vLLM writes 1 to the flag bit of the buffer to indicate the start of synchronization before starting the reduction summation, and resets the flag bit of the buffer to 0 after the reduction summation is completed. TensorRT increments the flag bit before each multi-card reduction summation to ensure that the flag bits used for each reduction summation are different. In addition, considering the case of CUDA Graph, when the graph capture is completed, the CUDA Kernel input parameters are fixed and will not be changed during actual calls. At this time, the method of incrementing the flag bit by TensorRT will fail. In response to this situation, the prior art needs to change the flag bit to a separate data pointer and increment the content inside the pointer each time. The above several operations all require additional processing of the flag bit, thus increasing the latency of the conversion between each multi-card reduction summation operation.

[0052] To this end, the present disclosure adopts a double-buffering mechanism of a double buffer (Double Buffer) to mark each reduction sum. That is, two buffers are created, denoted as Buffer1 and Buffer2. The flag bits corresponding to each buffer are Flag1 and Flag2 respectively. And a global counter (count) is maintained during the model inference process. Each time the multi-card reduction sum is called, the counter count is incremented by 1. A modulo 2 operation is performed on the counter to alternately use the two buffers. Since the flag bits used by each buffer are different, the purpose of distinguishing different multi-card reduction sums is achieved. Under the double-buffering mechanism, there is no need to reset the buffer or perform additional processing on the flag bits, achieving a lower latency.

[0053] In summary, according to the system and method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning of the present disclosure, considering that in the CUDA memory model, 64-bit atomic visibility is available, by combining 32-bit data and a 32-bit Flag bit into a 64-bit data packet each time, when the 32-bit Flag bit can be verified, it also means that the previous 32-bit data has been transmitted. Each thread in CUDA supports a maximum of 128-bit read and write, so two data packets can be written into the remote buffer at one time. The remote GPU reads the Buffer through busy wait and verifies whether it is the value of the Flag bit. When it is equal to the flag bit value, it indicates that the data has been successfully transmitted. The present disclosure utilizes the characteristics of 128Bytes sequential visibility of CUDA memory operations and the characteristics of each thread in CUDA supporting a maximum of 16Bytes read and write. Therefore, the data to be transmitted is taken as a data group with 128Bytes, and this data block is divided into 8 with 16Bytes as one thread, so that each group just meets the 128Bytes transmission limit of CUDA. For the last thread of each thread group, let it store 8Bytes of data and an 8Bytes Flag bit. This 8Bytes Flag bit is exactly 64 bits, meeting the 64-bit atomic visibility characteristic in the CUDA memory model, so that the data packet verification is only completed by this thread, and the other threads in the thread group all store 16Bytes of valid data. Therefore, when the Flag bit in the last thread of each thread group meets the expected value, it means that the previous 120Bytes of data have been successfully transmitted, which greatly utilizes the low-latency and high-bandwidth optimization protocol designed for collective communication between GPUs (such as All Reduce and Broadcast) in NVIDIA NCCL (cluster communication library), making the protocol bandwidth utilization rate reach more than 90%, and also improving the rate of multi-card reduction summation and reducing its latency from the perspective of bandwidth utilization. In addition, by reducing the data transmission volume during the reduction summation through two-stage reduction summation, the latency of the reduction summation is also reduced.

[0054] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from the present embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.

[0055] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0056] The above specifically shows and describes the exemplary embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the detailed structures, setting manners or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.

Claims

1. A system for reducing the latency of multi-GPU card reduction summation in distributed text reasoning, comprising: A data packing component that packs the data that needs to perform multi-GPU card reduction on the local GPU card according to a predetermined size, where the first part of each data packet is the data to be transmitted sequentially segmented from the shard data, and the second part is the flag bit data; A data transmission component that transmits the packed data packets from the local GPU card to the buffers of other remote GPU cards; And A data reception confirmation component that, during the polling wait process, when the reception of a data packet of a predetermined size is completed, polls for the completion of the corresponding flag bit data, and based on the visibility of data transmission, makes the staged reception completion result visible and confirmed to all GPU cards for multi-card reduction summation.

2. The system for reducing the latency of multi-GPU card reduction summation in distributed text reasoning according to claim 1, where the first part of each data packet is 32-bit data to be transmitted sequentially segmented from the shard data, the second part is 32-bit flag bit data, and the data transmission component transmits double data packets from the local GPU card to the buffers of other remote GPU cards each time. And the data reception confirmation component of each GPU card, during the polling wait process, when the reception of a 64-bit data packet is completed, polls for the completion of the corresponding flag bit data, and based on the 64-bit atomic visibility in the CUDA memory model, makes the staged reception completion result visible and confirmed to all GPU cards for multi-card reduction summation.

3. The system for reducing the latency of multi-GPU card reduction summation in distributed text reasoning according to claim 1, where the first part of each data packet is 120-byte data to be transmitted sequentially segmented from the shard data, the second part is 8-byte flag bit data, and the data transmission component transmits the data packets from the local GPU card to the buffers of other remote GPU cards based on the nvlink data transmission channel each time. And the data reception confirmation component of each GPU card performs a polling wait process to poll for the completion of the corresponding flag bit data when the reception of a 128-byte data packet is completed, and based on the sequential visibility in the CUDA memory model, makes the staged reception completion result visible and confirmed to all GPU cards for multi-card reduction summation.

4. The system for reducing the latency of multi-GPU card reduction summation in distributed text reasoning according to any one of claims 1-3, where the data transmission component of each GPU card transmits the continuously packed data packets of the local full amount of data from the local GPU card to the buffers of other remote GPU cards respectively, so that each GPU card performs multi-card reduction summation locally.

5. The system for reducing the latency of multi-GPU card reduction summation in distributed text reasoning as described in any one of claims 1-3, wherein the data packing component of each GPU card divides the local full amount of data into sub-slice data in an average manner corresponding to the number of GPU cards before packing, and then packs the sub-slice data respectively, and the data transmission component only transmits the consecutive data packets of the corresponding sub-slice data from the local GPU card to the buffer of the corresponding remote GPU card respectively, and after each GPU card receives all its corresponding sub-slice data and performs multi-card reduction summation on all the received sub-slice data locally to obtain the reduced summation result sub-slice data, each GPU card sends its local reduced summation result sub-slice data to other GPU cards.

6. A method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning, comprising: packing the data that needs to perform multi-GPU card reduction on the local GPU card by a data packing component according to a predetermined size, wherein the first part of each data packet is the data to be transmitted sequentially segmented from the sliced data, and the second part is the flag bit data; transmitting the packed data packets from the local GPU card to the buffer of other remote GPU cards by a data transmission component; and performing a polling waiting process by a data receiving and confirming component, so as to perform polling on the corresponding flag bit data when the reception of the data packets of the predetermined size is completed, and making the stage receiving completion result visible and confirmed to all GPU cards based on the visibility of data transmission, so as to perform multi-card reduction summation.

7. The method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning as described in claim 6, wherein the first part of each data packet is 32-bit data to be transmitted sequentially segmented from the sliced data, the second part is 32-bit flag bit data, and the data transmission component transmits double data packets from the local GPU card to the buffer of other remote GPU cards each time, and the data receiving and confirming component of each GPU card performs a polling waiting process, so as to perform polling on the corresponding flag bit data when the reception of 64-bit data packets is completed, and making the stage receiving completion result visible and confirmed to all GPU cards based on the 64-bit atomic visibility in the CUDA memory model, so as to perform multi-card reduction summation.

8. The method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning as claimed in claim 6, wherein the first part of each data packet is 120 bytes of data to be transmitted sequentially segmented from the sharded data, the second part is 8 bytes of flag bit data, and the data transmission component transfers the data packet from the local GPU card to the buffer of other remote GPU cards based on the nvlink data transmission channel each time, and the data reception confirmation component of each GPU card performs a polling waiting process so that when the reception of a 128-byte data packet is completed, the polling for the corresponding identification bit data is completed, and based on the sequential visibility in the CUDA memory model, the phased reception completion result is visible and confirmed by all GPU cards for multi-card reduction summation.

9. The method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning as claimed in any one of claims 6-8, wherein the data transmission component of each GPU card transfers the continuously packed data packets of the local full amount of data from the local GPU card to the buffers of other remote GPU cards respectively, so that each GPU card performs multi-card reduction summation locally.

10. The method for reducing the latency of multi-GPU card reduction summation in distributed text reasoning as claimed in any one of claims 6-8, wherein the data packing component of each GPU card evenly divides the sharded data into sub-sharded data corresponding to the number of GPU cards before packing, and then packs each sub-sharded data, and the data transmission component only transfers the continuously packed data packets of the corresponding sub-sharded data from the local GPU card to the buffers of the corresponding remote GPU cards respectively, and after each GPU card receives all its corresponding sub-sharded data and performs multi-card reduction summation on all the received sub-sharded data locally to obtain the reduced summation result sub-sharded data, it sends the local reduced summation result sub-sharded data to other GPU cards respectively.

Citation Information

Patent Citations

  • Instruction deferred execution and instruction specification method and device

    CN108280515A

  • Model training method, server and computer readable storage medium

    CN110134636A

  • Systems and methods for implementing network interface-based full reduction operations

    CN115686819A

  • Data transmission method and device, equipment and storage medium

    CN117215978A

  • Model deployment method and electronic equipment

    CN118313441A

Cited By

  • FlashMLA-based high-frequency transaction low-delay data processing system

    CN120950254A