Multi-GPU transformer parallel acceleration architecture and method based on distillation sampling

By introducing a distillation sampler into a multi-GPU transformer, sampling and transmitting data from a multi-GPU transformer, the problem of cross-GPU communication affecting computing efficiency is solved, and the inference speed is significantly improved.

CN120068996APending Publication Date: 2025-05-30ZHONGDIAN DATA IND CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411782780.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

There are a large number of cross-GPU communications in multi-GPU transformers, which affect computing efficiency.

Method used

Using a multi-GPU transformer parallel acceleration architecture based on distillation sampling, the data sampling is sampled by outputting results from the multi-headAttention layer of multiple GPUs under preset conditions, reducing the amount of data transmission across GPUs.

Benefits of technology

It significantly improves the inference speed of multi-GPUs, reduces the time-consuming memory access across GPUs, and the effect is not significantly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068996A_ABST
    Figure CN120068996A_ABST
Patent Text Reader

Abstract

The invention provides a multi-GPU (Graphics Processing Unit) transfer parallel acceleration architecture and a multi-GPU transfer parallel acceleration method based on distillation sampling. The multi-GPU transfer parallel acceleration architecture is characterized in that a sequence-head Attention layer of the multi-GPU transfer parallel architecture is connected with a Linear linear layer, and the sequence-head Attention layer of the multi-GPU transfer parallel architecture is connected with a concat layer; the multi-GPU transfer parallel acceleration architecture further comprises a distillation sampler, the distillation sampler is connected between the mutil-headAttention layer and the concat layer through a bypass, and the distillation sampler is used for outputting results from the mutil-headAttention layers of the multiple GPUs to carry out data sampling under a preset condition, obtaining sampling data and sending the sampling data to the concat layer. According to the method, the consumed time is optimized from the aspect of cross-GPU memory access, and the reasoning speed of the GPU is remarkably increased by reducing the cross-GPU data transmission quantity under the condition that the effect reduction is not obvious.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular, to a multi-GPU transformer parallel acceleration architecture and method based on distilled sampling. Background Art

[0002] Since 2017, the Transformer model has made significant progress, such as BERT and GPT. More recently, ChatGPT has achieved a high level of artificial intelligence. Although the performance of the Transformer is getting better and better, the number of parameters in the model is also increasing rapidly. A single GPU can no longer accommodate the model, and the inference speed is also getting slower and slower, which has greatly affected the application scenarios of the Transformer model. Therefore, many acceleration methods have emerged, such as model quantization technology, model tensor parallel inference, KV-cache, PageAttention, FlashAttention, etc. These methods have alleviated the inference performance problem to varying degrees.

[0003] For modern GPU inference, the main time-consuming occurs in the computational wall (large computational operation time) and the memory wall (large memory operation time). Moreover, in the case of multi-GPU, the memory wall problem is much more time-consuming than the computational time. Methods such as model quantization, multi-GPU tensor parallel acceleration, KV-cache, and PageAttention transparently accelerate the Transformer by optimizing the computation and VRAM, mainly optimizing the computational efficiency and memory loss. FlashAttention reduces memory operations at the cost of increased computation. However, there are still key time-consuming problems: that is, during multi-GPU inference, due to a large number of cross-GPU communications, the speed of these cross-GPU communications is very slow, resulting in the inability to fully utilize the computational efficiency of multi-GPU. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to solve the problem that a large number of cross-GPU communications in multi-GPU transformers affect the computational efficiency, and to provide a multi-GPU transformer parallel acceleration architecture and method based on distilled sampling.

[0005] According to the multi-GPU transformer parallel acceleration architecture based on distilled sampling in an embodiment of the present invention, a Linear layer is connected after the mutil-headAttention layer of the multi-GPU transformer parallel architecture, and a concat layer is connected after the Linear layer;

[0006] The multi-GPU transformer parallel acceleration architecture further includes: a distillation sampler, which is connected between the mutil-headAttention layer and the concat layer through a bypass. The distillation sampler is used to sample data from the output results of the mutil-headAttention layers of multiple GPUs under preset conditions, and send the sampled data to the concat layer.

[0007] According to some embodiments of the present invention, the concat layer supports setting probabilities to determine the probabilities of the concat layer using the complete data and sampled data of each GPU.

[0008] In some embodiments of the present invention, the distillation sampler performs sampling at the following positions:

[0009] (i + 0.5) * h / num;

[0010] where i = 0, 1, 2,....num, num is the number of samples, i is the i-th sampling point; h is the dimension of the multi-head layer.

[0011] According to some embodiments of the present invention, the distillation sampler supports setting a sampling probability to determine the number of samples num.

[0012] In some embodiments of the present invention, the distillation sampler samples training samples for model training. The input of the training sample X = (x, y_attention_statics), where x is the QKV input vector; y_attention_statics is the sampled data of y_attention and is expanded to the same length as y_attention, and y_attention is the data result of the mutil-headAttention layer.

[0013] According to the multi-GPU transformer parallel acceleration method based on distillation sampling of the embodiments of the present invention, the method uses the multi-GPU transformer parallel acceleration architecture based on distillation sampling as described above to accelerate and optimize the multi-GPU transformer. The method includes:

[0014] S10, load different multi-headAttention layers onto different GPUs, and load the corresponding distillation sampler according to the transformer layer index and the multi-headAttention layer index;

[0015] S20. Determine the probabilities for the Concat layer to select and use the complete data and sampled data of each GPU according to the probability P [0 to 1] set by the Concat layer; when P = 1, the Concat layer selects and uses the complete data of each GPU; when P < 1, the Concat layer selects and uses the sampled data of each GPU.

[0016] According to some embodiments of the present invention, in step S20, when the Concat layer selects and uses the sampled data of each GPU, the distillation sampler determines the sampling amount according to the set sampling probability, and samples to obtain the sampled data and sends it to the Concat layer.

[0017] In some embodiments of the present invention, before executing the multi-GPU transformer parallel acceleration method, the distillation sampler is trained.

[0018] The present invention has the following beneficial effects:

[0019] The present invention optimizes the time-consuming aspect from cross-GPU memory access. Without an obvious reduction in the effect, by reducing the cross-GPU data transfer volume, the inference speed of multiple GPUs is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram of a traditional multi-GPU transformer architecture;

[0021] Figure 2 is a schematic diagram of a multi-GPU transformer parallel acceleration architecture based on distillation sampling according to an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of a GPU data synchronization strategy using a ring. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention is described in detail as follows in combination with the accompanying drawings and preferred embodiments.

[0024] In the present invention, the description of the method process in the specification and the steps in the flowchart in the accompanying drawings of the present invention do not necessarily have to be strictly executed according to the step numbers. The execution order of the method steps can be changed. Moreover, certain steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0025] The present invention innovatively designs a multi-GPU transformer parallel acceleration architecture and method based on distilled sampling to reduce the communication frequency across GPUs and improve the parallel inference efficiency of multi-GPUs. First, the transformer model architecture is adjusted based on matrix decomposition, and the concat timing of the multi-head is postponed. Then, the multi-head layer of the transformer is sampled using neural network distillation, and the cross-GPU communication is controllably reduced during the ring cycle to achieve the low-loss acceleration effect of the model.

[0026] The present invention innovatively proposes a multi-GPU transformer parallel acceleration architecture based on distilled sampling, which reduces the cross-GPU communication volume and significantly improves the inference speed of multi-GPUs when the model effect does not decrease significantly.

[0027] As Figure 2 shown, in the multi-GPU transformer parallel acceleration architecture based on distilled sampling according to an embodiment of the present invention, a Linear layer is connected after the mutil-headAttention layer of the multi-GPU transformer parallel architecture, and a concat layer is connected after the Linear layer;

[0028] Compare Figure 1 and Figure 2 It can be seen that in the present invention, the Linear layer is moved forward to before the concat layer. It should be noted that the Linear layer is a linear operation. According to the matrix decomposition rule, the matrix operation can be implemented in blocks to achieve the decomposition effect. Thus, the data interaction volume between GPUs can be reduced, the computing efficiency of the multi-GPU transformer can be improved, and the equivalence of the results can be maintained.

[0029] For example, assume that there are a total of 4 GPUs. Only two parameter matrices, A1 (the output of the Attention layer on this GPU) and B (the parameters of the linear layer), are on one GPU, and parts such as A2, A3, and A4 are on other GPUs. According to the standard algorithm, it should be: first collect the A parts on each GPU, concatenate them together (concat), and then pass them into the Linear layer to calculate A*B.

[0030] Let the original matrix be A m×n , B n×p , then the original calculation method is:

[0031]

[0032] After the Linear layer of the present invention is placed in front, on the GPU with A1, only the content of A1*B1 is calculated. Similarly, other GPUs calculate A2*B2, A3*B3, and A4*B4. After each of the four GPUs has completed its calculation, the results are passed into the concat layer for splicing. At this time, the accuracy of the results can still be maintained, and the aggregation of the outputs of the GPUAttention layers of each GPU is avoided, thereby reducing the communication volume between GPUs and improving the calculation efficiency of the model.

[0033] The multi-GPU transformer parallel acceleration architecture further includes: a distillation sampler, which is connected between the mutil-headAttention layer and the concat layer through a bypass. The distillation sampler is used to sample data from the output results of the mutil-headAttention layers of multiple GPUs under preset conditions, and send the sampled data to the concat layer.

[0034] It should be noted that, as Figure 3 shown, a variable data is distributed and stored in N GPUs. Each GPU stores 1 / N of it. If a GPU wants to obtain all the data, it needs to copy the data of the other N-1 GPUs. To improve the throughput of copying data, the ring multi-GPU synchronization strategy is often used now. For example Figure 3 shown, the ring moves N-1 times, and each GPU can obtain the complete data of the variable.

[0035] In the above process, each GPU needs to transmit its own complete data each time, and the cross-GPU data transmission bandwidth is low and time-consuming is high. The present invention adopts a distillation sampler to effectively reduce the cross-GPU data transmission.

[0036] According to some embodiments of the present invention, the concat layer supports setting probabilities to determine the probabilities of the concat layer selecting to use the complete data and sampled data of each GPU.

[0037] As Figure 2 shown, there are two data sending methods for concat in the present invention:

[0038] Transmit the owned target data completely to other GPUs;

[0039] Transmit the owned target data to other GPUs according to the sampling method, and use the sampling information of the target data here;

[0040] Concat has two data receiving methods:

[0041] Receive the complete data transmitted by other GPUs;

[0042] Receive the sampled data passed from other GPUs, and then call the distilled sampler to generate sampled data;

[0043] The probability P [0-1] set by the Concat layer determines the probability that the Concat layer operator uses the complete target data and sampled data. The Concat operator automatically determines whether to use the distilled sampler according to the received data format. The main purpose of the distilled sampler is to indirectly fit the data distribution of the multi-head on other GPUs, thereby reducing the cross-GPU data transmission.

[0044] It should be noted that a decorator can be constructed by rewriting the contact operator to map the Concat operator behind the multi-head in the original transformer, so as to achieve a user-transparent model replacement.

[0045] In some embodiments of the present invention, the sampling algorithm fun_t of the distilled sampler performs sampling at the following positions:

[0046] (i + 0.5)*h / num;

[0047] where i = 0, 1, 2,....num, num is the number of samples, i is the i-th sampling point; h is the dimension of the multi-head layer, such as 4096. After sampling, a sampling result matrix is obtained.

[0048] According to some embodiments of the present invention, the distilled sampler supports setting a sampling probability to determine the number of samples num. For example, if the sampling probability T is determined to be between [0, 100%], the number of samples is calculated as: num = int(h*T).

[0049] Generally speaking, by adjusting T and P, the balance between the model inference accuracy and the inference speed can be dynamically regulated, which is convenient for users to select different balance points according to different scenarios.

[0050] In some embodiments of the present invention, the distilled sampler samples training samples for model training. The input X of the training sample is (x, y_attention_statics), where x is the QKV input vector; y_attention_statics is the sampled data of y_attention, and is expanded into a diff format with the same length as y_attention. The expansion method can be to supplement 0. y_attention is the data result of the mutil-head Attention layer. The input X of the training sample is (x, y_attention_statics).

[0051] For example, a pre-set question data set can be used to save the input QKV and output results of different multi-head attention layers, as well as the output statistical sampling data to train the distillation sampler for the model, including:

[0052] The distillation model F uses a linear model and a direction matrix with the same matrix dimension as the X dimension;

[0053] Training method: y_attention = F(X);

[0054] Loss function: mean square error function, which measures the fitting ability of the distillation model;

[0055] Finally, save the distillation sampling model to disk.

[0056] Each GPU stores Num_transformer * (Num_multi-head - 1) distillation sampling models;

[0057] Among them, Num_transforme is the number of Num_transformer layers stored in the GPU, and Num_multi-head is the number of multi-heads in the corresponding transform;

[0058] According to the multi-GPU transformer parallel acceleration method based on distillation sampling of the embodiments of the present invention, the method uses the above-mentioned multi-GPU transformer parallel acceleration architecture based on distillation sampling to accelerate and optimize the multi-GPU transformer. The method includes:

[0059] S10, based on the tensor parallel strategy, when loading the standard transformer model, load different multi-head attention layers onto different GPUs, and load the corresponding distillation sampler according to the transformer layer index and multi-head attention layer index;

[0060] S20, determine the probability of the concat layer to select and use the complete data and sampling data of each GPU according to the probability P [0-1] set by the Concat layer; when P = 1, the concat layer selects and uses the complete data of each GPU; when P < 1, the concat layer selects and uses the sampling data of each GPU.

[0061] According to some embodiments of the present invention, in step S20, when the concat layer selects and uses the sampling data of each GPU, the distillation sampler determines the sampling amount according to the set sampling probability, and samples to obtain the sampling data and sends it to the concat layer.

[0062] In some embodiments of the present invention, before performing the multi-GPU transformer parallel acceleration method, model training is performed on the distilled sampler.

[0063] In summary, the improved multi-GPU tensor parallel transformer acceleration architecture and method of the present invention utilize matrix decomposition to advance the linear layer of the transform and delay the multi-head concat timing to reduce sampling errors.

[0064] Train the distilled sampling model for different multi-heads (each GPU has a part of the multi-head and the distilled sampling model of another part of the multi-head during inference) (the sampling algorithm can control the sampling accuracy through the parameter T).

[0065] During the concat operation, fully synchronize the data of part of the GPUs for the probability P, and only synchronize the sampled data for other GPUs. Use the distilled sampling model and the sampled data of other GPUs to sample the data of the multi-head where it is located.

[0066] The present invention has the following beneficial effects:

[0067] The present invention optimizes the time-consuming aspect from cross-GPU memory access (cross-GPU operations are significantly slower than calculations and memory operations). When the effect reduction is not obvious, it can significantly improve the inference speed of multi-GPUs (by reducing the cross-GPU data transfer volume).

[0068] Through the description of the specific implementation manner, it should be possible to have a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose. However, the accompanying drawings are only for reference and illustration purposes and are not used to limit the present invention.

Claims

1. A multi-GPU transformer parallel acceleration architecture based on distillation sampling, characterized in that: The mutil-headAttention layer of the multi-GPU transformer parallel architecture is connected to a Linear layer, and the Linear layer is connected to a concat layer; The multi-GPU transformer parallel acceleration architecture also includes: a distillation sampler, which is connected between the mutil-headAttention layer and the concat layer through a bypass, and the distillation sampler is used to sample data from the output results of the mutil-headAttention layers of multiple GPUs under preset conditions, and obtain the sampled data and send it to the concat layer.

2. The multi-GPU transformer parallel acceleration architecture based on distillation sampling according to claim 1, characterized in that: The concat layer supports setting probabilities to determine the probabilities of the concat layer choosing to use complete data and sampled data of each GPU.

3. The multi-GPU transformer parallel acceleration architecture based on distillation sampling according to claim 1, characterized in that: The distillation sampler performs sampling at the following locations: (i+0.5)*h / num; Among them, i=0, 1, 2, ....num, num is the sampling amount, i is the i-th sampling point; h is the dimension of the multi-head layer.

4. The multi-GPU transformer parallel acceleration architecture based on distillation sampling according to claim 3, characterized in that: The distillation sampler supports setting the sampling probability to determine the sampling amount num.

5. The multi-GPU transformer parallel acceleration architecture based on distillation sampling according to claim 1, characterized in that: The distillation sampler samples training samples for model training. The input of the training samples is X=(x, y_attention_statics), where x is the QKV input vector; y_attention_statics is the sampled data of y_attention and is expanded to the same length as y_attention. y_attention is the data result of the mutil-headAttention layer.

6. A multi-GPU transformer parallel acceleration method based on distillation sampling, characterized in that: The method uses a multi-GPU transformer parallel acceleration architecture based on distillation sampling as described in any one of claims 1-5 to accelerate and optimize the multi-GPU transformer, and the method includes: S10, load different multi-headAttention layers onto different GPUs, and load the corresponding distillation samplers according to the transformer layer index and the multi-headAttention layer index; S20, determine the probability of the concat layer choosing to use the complete data and sampled data of each GPU according to the probability P[0~1] set by the Concat layer; when P=1, the concat layer chooses to use the complete data of each GPU; when P<1, the concat layer chooses to use the sampled data of each GPU.

7. The multi-GPU transformer parallel acceleration method based on distillation sampling according to claim 6, characterized in that: In step S20, when the concat layer chooses to use the sampling data of each GPU, the distillation sampler determines the sampling amount according to the set sampling probability, and samples to obtain the sampling data and sends it to the concat layer.

8. The multi-GPU transformer parallel acceleration method based on distillation sampling according to claim 6, characterized in that: The distilled sampler is trained before executing the multi-GPU transformer parallel acceleration method.