A method, system and device for Transformer model inference calculation
By dividing the calculation graph of the Transformer model into a structure in which the calculation sub-graph and attention operator are alternately connected, and using different types of computing devices for inference calculation, the problems of low resource utilization and high cost in the existing technology are solved, and the calculation speed and efficiency are improved.
Patent Information
- Application Number
- CN202311482067.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2043-11-08
AI Technical Summary
In the inference calculation of existing Transformer models, computing device resource utilization is low, inefficient and costly, especially for large neural networks and large language models.
The calculation diagram of the Transformer model is divided into a structure in which the calculation sub-graph and attention operator are alternately connected, and different types of computing devices are used for inference calculation, including type A and type B computing devices. Type A devices are stronger than type B devices' computing capabilities, and type B devices have large storage space and bandwidth.
It significantly improves hardware resource utilization, improves computing speed and reduces costs, especially for large models, the inference computing effect is significant.
Smart Images

Figure CN117725963B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a computing method, system, and device for Transformer model reasoning. Background Art
[0002] Currently, most large models are built based on the Transformer model. The Transformer model is a machine learning model that incorporates an attention mechanism and a feedforward neural network. Transformer model inference refers to the process of obtaining user input data, feeding that data into the inference system for calculation, and obtaining the calculated results.
[0003] The current common approach for reasoning with larger Transformer models is to use one or more computing devices of the same type. Specifically, when the model is relatively small, a single computing device is used to complete the entire model inference process. When the model is relatively large, the entire model is divided into several parts (Part 1, Part 2, and so on) to be processed serially. The model input and Part 1 are processed by the first computing device, while the results of Part 1 and Part 2 are processed by the second computing device of the same type. This continues until the results of Part 1, Part 2, are obtained as the final results of the entire model.
[0004] Existing Transformer models all use the same type of computing equipment to complete the entire inference process. However, different parts of the Transformer model have different performance requirements for different computing equipment. Therefore, using only a single type of computing equipment for Transformer model inference, especially for inference calculations on large-scale Transformer models (such as large neural networks and large language models), makes it difficult to achieve high resource utilization of the computing equipment on a sustained basis. This leads to low inference efficiency and high overall cost.
[0005] There is currently no effective method to improve computing efficiency and resource utilization of computing devices and reduce computing costs in Transformer model inference calculations. Summary of the Invention
[0006] The purpose of the present invention is to solve the problems of low reasoning efficiency, low computing device resource utilization and high cost of existing Transformer models. A method, system and device for Transformer model reasoning calculation are proposed, which divides the calculation graph of a Transformer model into a structure consisting of alternating calculation subgraphs and attention operators, wherein the calculation subgraph part is calculated using the first type of computing device, and the attention operator is calculated using the second type of device.
[0007] The present invention is achieved through the following solutions:
[0008] A method for inference calculation of a Transformer model, comprising: dividing a computation graph of a Transformer model into a structure consisting of computation subgraphs and attention operators alternately connected in series; and performing inference calculation using two computing devices of different models or types.
[0009] Furthermore, when the input data scale of the attention operator increases, the growth rate of its computational complexity will not exceed the growth rate of the input data scale.
[0010] Furthermore, the structure of alternately connecting attention operators and computation subgraphs includes n+1 computation subgraphs and n attention operators, where n is a natural number; the computation devices are type A computation devices and type B computation devices; and the inference computation process includes:
[0011] Merge and copy the input data of multiple groups of inference requests to the A-type computing device to perform calculations. Figure 1 The inference calculation of Figure 1 The inference calculation results of the attention operator 1 are split into several parts, copied to the corresponding number of B-type computing devices, and the inference calculation of the attention operator 1 is performed to obtain the calculation results; the output results of the attention operator 1 corresponding to multiple requests are merged and copied to the A-type computing device, and the calculation results are performed. Figure 2 The inference calculation of subgraph n+1 is performed and the calculation result is obtained; and so on, until the inference calculation result of subgraph n+1 is obtained; the inference calculation result of subgraph n+1 is split into a single inference result as the result of inference calculation on the entire model corresponding to the inference request.
[0012] Furthermore, the total computing power of the type A computing device is stronger than that of the type B computing device; and the total storage space of the memory of the type B computing device is greater than that of the memory of the type A computing device.
[0013] Furthermore, the type A computing device and the type B computing device include: a CPU, a GPU, a TPU, an FPGA, an ASIC, a memory module with computing functions, and a whole formed by the CPU, GPU, TPU, FPGA, ASIC, and a memory module with computing functions interconnected through a high-speed bus / network.
[0014] The present invention also provides a system for Transformer model inference calculation, which is composed of a model segmentation module, a scheduling module, a type A inference calculation module, and a type B inference calculation module; the model segmentation module divides the calculation graph of the Transformer model into a structure consisting of attention operators and calculation subgraphs alternately connected in series; according to the result of the model segmentation, an inference system is constructed by connecting the scheduling module, the type A inference calculation module, and the type B inference calculation module.
[0015] Furthermore, the type A reasoning and computing module includes a type A computing device, and the type B reasoning and computing module includes a type B computing device.
[0016] Furthermore, each of the computational subgraphs has a corresponding scheduling module.
[0017] Furthermore, the calculator Figure 1 The corresponding scheduling module 1 is responsible for collecting the reasoning requests of the entire Transformer model. Figure 1 When the corresponding A-type computing device is idle, the input of the inference request is merged and copied to the A-type computing device, and the corresponding A-type inference computing module is notified to drive the A-type device to perform the calculation. Figure 1 calculation; calculator Figure 1 After the calculation is completed, the A-type device calculation module splits the calculation results and waits for the B-type computing device corresponding to the attention operator 1 to be idle, copies the calculation results to the B-type computing device, and notifies the corresponding B-type reasoning calculation module to drive the B-type computing device to perform calculations; when the B-type reasoning calculation module receives the notification, it calculates the attention operator 1 and then sends the calculation results to the calculation device. Figure 2 Scheduling module 2; and so on, until the A-type reasoning calculation module of the calculation subgraph n+1 obtains the calculation result, the calculation result is split to obtain the reasoning result corresponding to each reasoning request.
[0018] The present invention further provides a device for Transformer model inference calculation, which includes a memory, a processor, and a computer program stored on the memory and runnable on the processor; when the processor executes the computer program, the aforementioned method for Transformer model inference calculation and system for Transformer model inference calculation are implemented.
[0019] By segmenting the Transformer model and using different types of computing devices for different inference calculations, this approach significantly improves hardware resource utilization, increases computing speed, and reduces inference costs compared to conventional methods that rely on the same type of devices. This approach is particularly effective for inference on large models (e.g., large neural networks and large language models). BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of the segmentation of the computational graph of the Transformer model in the present invention is shown.
[0022] Figure 2 A schematic diagram of a process for Transformer model inference calculation according to an embodiment of the present invention is shown;
[0023] Figure 3 A system block diagram for Transformer model inference calculation provided by the present invention is shown;
[0024] Figure 4 A schematic diagram of the partitioning of a Transformer model consisting of n Transformer blocks is shown. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0026] The present invention provides a method for inference calculation of a Transformer model. First, a calculation graph of a Transformer model is divided into a structure consisting of calculation subgraphs and attention operators alternately connected in series, as shown in the attached figure. Figure 1 As shown. Among them, the attention operator is a
[0027]
[0028] The operator of the calculation part, where Q is the query vector or the query matrix composed of multiple query vectors, K and V are the key matrix and value matrix, d k is the dimension of the vector.
[0029] It's important to note that the attention operator is a type of machine learning operator that dynamically assigns weights to a set of input vectors and combines them to produce an output. While there are many variations and improvements to the attention operator, they all share the following characteristic: their computational complexity is no more than linear. This means that as the size of the operator's input data increases, the computational effort does not grow faster than the size of the input data.
[0030] Then, in a preferred embodiment, two different models or types of computing devices, namely type A computing device and type B computing device, are used to perform the inference calculation process. Figure 2 shown.
[0031] See attached Figure 1 and attached Figure 2 , the reasoning calculation process of the Transformer model after segmentation is as follows:
[0032] (1) Merge and copy the input data of multiple groups of inference requests to the A-type computing device to perform calculations. Figure 1 Perform reasoning calculations and obtain calculation results;
[0033] (2) The calculator Figure 1 The inference calculation result of is split into several parts, which are copied to a number of corresponding B-type computing devices and the inference calculation of attention operator 1 is performed to obtain the calculation result;
[0034] Of course, those skilled in the art will understand that the calculation Figure 1 The inference calculation results can also be directly output to a Type B computing device without splitting, and then the inference calculation of attention operator 1 is performed to obtain the calculation results. In other words, the focus of the present invention is to use different types of computing devices to perform inference calculations on the segmentation results of the Transformer model, and select computing devices with appropriate performance indicators according to the specific needs of data calculation to improve the inference calculation speed.
[0035] (3) Merge the output results of attention operator 1 corresponding to multiple requests and copy them to type A computing device for calculation. Figure 2 Perform reasoning calculations and obtain calculation results;
[0036] (4) This process is repeated until the inference calculation result of computation subgraph n+1 is obtained. The inference calculation result of computation subgraph n+1 is split into individual inference results, which are used as the inference calculation result of the corresponding inference request on the entire model.
[0037] Generally speaking, when the Transformer model performs inference, the attention operator has high requirements for the storage capacity and bandwidth of the computing device, while other parts such as the computational subgraph have high requirements for the computing power of the computing device. In order to improve device utilization and reduce computing time, the type A computing device in the above-mentioned optimal embodiment usually adopts a device with stronger computing power, and the type B computing device usually adopts a device with larger storage space and larger bandwidth memory. In other words, the total computing power of type A computing devices (the number is greater than or equal to 1) is stronger than that of type B computing devices; the total storage space of type B computing devices (the number is greater than or equal to 1) is greater than that of type A computing devices.
[0038] The aforementioned Type A and Type B computing devices include, but are not limited to, CPUs (central processing units), GPUs (graphics processing units), TPUs (tensor processing units), FPGAs (field programmable gate arrays), ASICs (application-specific integrated circuits), memory modules with computing capabilities, and their interconnection via high-speed buses or networks. Memory modules with computing capabilities typically employ a technology known as "processing-in-memory" (PIM) or "near-memory processing." Their fundamental characteristic is the addition of computing circuitry within a memory chip or the addition of interconnected computing chips near the memory chip. This allows for a level of computing power comparable to conventional memory devices (which only provide read and write functions). Such products are particularly suitable for use as the Type B computing devices of the present invention.
[0039] As attached Figure 3 As shown, the present invention also provides a system for Transformer model inference calculation. The system consists of a model segmentation module, a scheduling module, a type A inference calculation module, and a type B inference calculation module.
[0040] The model splitting module splits the computational graph of the Transformer model into a structure consisting of attention operators and computational subgraphs alternately connected in series.
[0041] Then, based on the results of model segmentation, an inference system is constructed which is connected by a scheduling module, a type A inference calculation module, and a type B inference calculation module.
[0042] The type A reasoning computing module includes a type A computing device, and the type B reasoning computing module includes a type B computing device.
[0043] See attached Figure 3 , each computational subgraph has a corresponding scheduling module. Figure 1 The corresponding scheduling module 1 is responsible for collecting the reasoning requests of the entire Transformer model (request 1, request 2, ... request k). When the preset conditions are met (the number of requests exceeds a certain number or a preset time has passed) and the calculation submodule Figure 1 When the corresponding A-type computing device is idle, the input of the inference request is merged and copied to the A-type computing device, and the corresponding A-type inference computing module is notified to drive the A-type device to perform the calculation. Figure 1 calculation; calculator Figure 1 After the calculation is completed, the A-type device calculation module splits the calculation results and waits for the B-type computing device corresponding to the attention operator 1 to be idle, copies the calculation results to the B-type computing device, and notifies the corresponding B-type reasoning calculation module to drive the B-type computing device to perform calculations; when the B-type reasoning calculation module receives the notification, it calculates the attention operator 1 and then sends the calculation results to the calculation device. Figure 2 Scheduling module 2. And so on, until the A-type reasoning calculation module of the calculation subgraph n+1 obtains the calculation result, the calculation result is split to obtain the reasoning results corresponding to each reasoning request (request 1, request 2, ... request k).
[0044] Attachment Figure 3 In the illustrated embodiment, Type A computing devices are used for inference calculations of the computational subgraph, and Type B computing devices are used for inference calculations of the attention operator. However, in specific application scenarios, those skilled in the art can flexibly select computing devices based on the actual device feature requirements for data computation.
[0045] In addition, the present invention also provides a device for Transformer model inference calculation, which includes a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the processor executes the computer program, it implements the method and system for Transformer model inference calculation as described above.
[0046] Example 1
[0047] Attachment Figure 4 A schematic diagram of the partitioning of a Transformer model consisting of n Transformer blocks is shown.
[0048] For an exemplary Transformer model consisting of n Transformer blocks, each Transformer block is composed of a first linear transformation operator, an attention operator, a second linear transformation operator, and a feedforward neural network. For the Transformer model of this embodiment, the process of performing inference calculation using the present invention is as follows.
[0049] The Transformer model of the above embodiment is segmented using the aforementioned segmentation module to obtain the following segmentation results:
[0050] The first computational subgraph includes the first linear transformation operator in Transformer block 1;
[0051] The first attention operator is the attention operator in Transformer block 1;
[0052] The second computational subgraph includes the second linear transformation operator in the Transformer1 block, the feedforward neural network operator, and the first linear transformation operator in the Transformer block 2;
[0053] The second attention operator is the attention operator in Transformer block 2;
[0054] …
[0055] The nth computational subgraph includes the second linear transformation operator in Transformer block n-1, the feedforward neural network operator, and the first linear transformation operator in Transformer block n;
[0056] The nth attention operator is the attention operator in Transformer block n;
[0057] The n+1th computational subgraph includes the second linear transformation operator and the feedforward neural network operator in Transformer block n.
[0058] Based on the above segmentation results, the GPU is used as the Type A computing device, and the CPU and its connected RAM are used as the Type B computing device. An inference system is constructed consisting of n scheduling modules, n+1 Type A inference computing modules, and n groups of Type B inference computing modules. Each Type B inference computing module group contains m Type B inference computing modules. Each Type A inference computing module uses one Type A technical device for calculation, and each Type B inference computing module uses one Type B design computing device for calculation.
[0059] The reasoning process of the inference computing system is as follows:
[0060] First, the scheduling module 1 corresponding to the first linear transformation operator collects inference requests. When the number of requests reaches a preset value (such as 100) and the type A computing device corresponding to the first type A inference calculation module is idle, all inputs are merged and copied to the type A device corresponding to the first type A inference calculation module, and the first type A inference calculation module is notified.
[0061] After receiving the notification, the first A-type reasoning calculation module performs the calculation Figure 1 , that is, the calculation of the first linear transformation operator in Transformer block 1, and split the result into m parts, respectively copy them to the 1st to mth B-type computing devices in the first group of B-type reasoning computing modules, and notify the corresponding B-type reasoning computing modules respectively.
[0062] After receiving the notification, the first group of type B reasoning modules will calculate the attention operator in Transformer block 1 on the received input and send the result to the second scheduling module.
[0063] The execution process of the second scheduling module is the same as that of the first scheduling module.
[0064] The execution process of the second A-type reasoning calculation module refers to the first A-type reasoning calculation module, but the calculation content is the calculation submodule. Figure 2 , namely the second linear transformation operator in Transformer block 1, the feedforward neural network operator, and the first linear transformation operator in Transformer block 2.
[0065] This process is deduced in this way until the (n+1)th type A inference calculation module obtains the calculation result. After splitting, the inference results corresponding to each inference request are obtained.
[0066] During the inference calculation process, different computing devices can perform inference calculations in parallel, thereby improving the efficiency of inference calculations.
[0067] By using different types of computing devices for different computing parts, we can fully utilize the advantages of different types of computing devices, improve the resource utilization of computing devices, improve the efficiency of reasoning calculations, and reduce costs.
[0068] Example 2
[0069] Same as Example 1, but there is only one module in each group of Type B reasoning modules. The process of sending the calculation results of the calculation subgraph by the Type A reasoning module to the subsequent Type B calculation module is not split, and the process of sending the results by the Type B module to the subsequent scheduling module does not need to be merged.
[0070] Example 3
[0071] Same as Example 1, but all Type A reasoning modules share one Type A computing device, and all Type B reasoning modules share one Type B computing device.
[0072] Example 4
[0073] Same as Example 1, but using TPU as type A computing device and memory module with computing function as type B computing device.
[0074] Example 5
[0075] Same as Example 1, but the example model has a preprocessing operator before the Transformer block 1. When splitting the Transformer model, the preprocessing operator is placed before the operator. Figure 1 at the front.
[0076] By segmenting the Transformer model and using different types of computing devices for different inference calculations, this approach significantly improves hardware resource utilization, increases computing speed, and reduces inference costs compared to conventional methods that rely on the same type of devices. This approach is particularly effective for inference on large models (e.g., large neural networks and large language models).
[0077] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for Transformer model inference calculation, characterized in that: The method comprises: The computational graph of a Transformer model is divided into a structure consisting of alternating computational subgraphs and attention operators; The process of performing inference calculations using two different models or types of computing devices; The computing devices are type A computing devices and type B computing devices; the total computing power of the type A computing devices is greater than that of the type B computing devices; and the total storage space of the memories of the type B computing devices is greater than that of the memories of the type A computing devices; The computation subgraph is inferred and calculated by a type A computing device; the attention operator is inferred and calculated by a type B computing device.
2. The method according to claim 1, characterized in that When the input data scale of the attention operator increases, the growth rate of its computation amount does not exceed the growth rate of the input data scale.
3. The method according to claim 1 or 2, characterized in that The structure of alternately connecting computation subgraphs and attention operators comprises n+1 computation subgraphs and n attention operators, where n is a natural number. The process of inference calculation includes: Merge and copy the input data of multiple groups of inference requests to the type A computing device, perform the inference calculation of computation subgraph 1, and obtain the calculation results; Split the inference calculation results of computation subgraph 1 into several parts, copy them to a corresponding number of B-type computing devices, and perform inference calculations on attention operator 1 to obtain the calculation results; Merge the output results of attention operator 1 corresponding to multiple requests and copy them to type A computing device to perform inference calculation of computation subgraph 2 and obtain the calculation results; This process is repeated until the inference calculation result of computation subgraph n+1 is obtained. The inference calculation result of computation subgraph n+1 is split into individual inference results, which serve as the inference calculation result of the corresponding inference request on the entire model.
4. The method according to claim 3, characterized in that The type A computing device and type B computing device include: CPU, GPU, TPU, FPGA, ASIC, memory modules with computing functions, and a whole formed by the CPU, GPU, TPU, FPGA, ASIC, and memory modules with computing functions interconnected by a high-speed bus / network.
5. A system for Transformer model inference calculation, used to execute the method for Transformer model inference calculation according to any one of claims 1 to 4, characterized in that: The system consists of a model segmentation module, a scheduling module, a type A reasoning calculation module, and a type B reasoning calculation module; The model splitting module splits the computational graph of the Transformer model into a structure consisting of attention operators and computational subgraphs alternately connected in series; According to the results of model segmentation, an inference system is constructed which is connected by a scheduling module, a type A inference calculation module, and a type B inference calculation module.
6. The system according to claim 5, characterized in that The type A reasoning and computing module includes a type A computing device, and the type B reasoning and computing module includes a type B computing device.
7. The system according to claim 5 or 6, characterized in that Each of the computational subgraphs has a corresponding scheduling module.
8. The system according to claim 7, characterized in that Scheduling module 1, corresponding to computational subgraph 1, is responsible for collecting inference requests for the entire Transformer model. When the preset conditions are met and the type A computing device corresponding to computational subgraph 1 is idle, it merges and copies the input of the inference requests to the type A computing device and notifies the corresponding type A inference computing module to drive the type A device to perform computations on computational subgraph 1. After the calculation of subgraph 1 is completed, the A-type device calculation module splits the calculation results and waits for the B-type computing device corresponding to attention operator 1 to be idle. Then, the calculation results are copied to the B-type computing device and the corresponding B-type inference calculation module is notified to drive the B-type computing device to perform calculations. When the B-type reasoning calculation module receives the notification, it calculates the attention operator 1 and then sends the calculation result to the scheduling module 2 of the calculation subgraph 2; The process continues in this way until the A-type reasoning calculation module of the calculation subgraph n+1 obtains the calculation result, and then the calculation result is split to obtain the reasoning results corresponding to each reasoning request.
9. A device for Transformer model inference calculation, characterized in that: The apparatus comprises a memory, a processor, and a computer program stored on the memory and executable on the processor; When the processor executes the computer program, it implements the method for Transformer model reasoning calculation as described in any one of claims 1-4 and the system for Transformer model reasoning calculation as described in any one of claims 5-8.
Citation Information
Patent Citations
Neural network model reasoning method and device, equipment and storage medium
CN116629308A
Image processing method and device, equipment and storage medium
CN116958760A