An operator execution method, device, storage medium and program product
By allocating tensor computing cores and vector computing cores on the artificial intelligence chip, and using multi-threaded bundle groups to process batch tasks in parallel, the problem of high time overhead for Attention operators when processing multiple batch tasks is solved, achieving more efficient processing efficiency.
Patent Information
- Application Number
- CN202510031673.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
When large language models perform inference on multiple batch tasks, the time overhead caused by matrix calculation and vector calculation of the Attention operator is too high. How to reduce the time overhead of inference on a large batch of tasks is an urgent problem.
By allocating tensor computing cores and vector computing cores on the artificial intelligence chip, multi-threaded bundles are used to process batch tasks in parallel. Specifically, the second thread bundle group uses vector operators to perform vector calculations for a certain batch of tasks, while the first thread bundle group and the third thread bundle group use matrix operators to perform matrix calculations for adjacent batches of tasks, making full use of the processing capabilities of the tensor calculation unit and the vector calculation unit.
Through this method, the time overhead brought by the Attention operator to process multiple batch tasks can be effectively reduced, and processing efficiency can be improved, while reducing the complexity of managing thread bundle groups.
Smart Images

Figure CN119416826B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to an operator execution method, device, storage medium, and program product. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) are widely used in the fields of natural language processing, computer vision, speech recognition, etc. Currently, large language models generally use the Transformer structure as the core infrastructure, and the core computing part of the Transformer structure is the attention mechanism (Attention) operator.
[0003] When the large language model performs inference on multiple batches of tasks, the Attention operator can sequentially perform matrix calculations and vector calculations for each batch of tasks. However, when there are too many batches of tasks to be inferred, this will bring huge time overhead. How to reduce the time overhead brought by inferring a large number of batches of tasks is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0004] Embodiments of the present application provide an operator execution method, device, storage medium, and program product for reducing the time overhead brought by inferring a large number of batches of tasks.
[0005] On the one hand, embodiments of the present application provide an operator execution method, which is applied to an artificial intelligence chip including S tensor calculation cores and W vector calculation cores. The artificial intelligence chip is used to process N batches of tasks by using the attention mechanism operator; the attention mechanism operator at least includes a first matrix operator, a vector operator, and a second matrix operator that are sequentially executed for any batch of tasks among the N batches of tasks; S and W are both positive integers, and N is an integer greater than 2;
[0006] When the vector operator is used in the second thread block group to perform vector calculations on the nth batch of tasks among the N batches of tasks, the first thread block group uses the first matrix operator to perform matrix calculations on the (n + 1)th batch of tasks among the N batches of tasks, and the third thread block group uses the second matrix operator to perform matrix calculations on the (n - 1)th batch of tasks among the N batches of tasks; the second thread block group includes the W vector calculation cores, the first thread block group includes S1 tensor calculation cores among the S tensor calculation cores, the third thread block group includes S2 tensor calculation cores among the S tensor calculation cores, and there is no overlap between the S1 tensor calculation cores and the S2 tensor calculation cores; the value range of n is [2, N], both S1 and S2 are less than S, and the sum of S1 and S2 is less than or equal to S.
[0007] Optionally, the second thread bundle group performs vector calculation on the nth batch of tasks among the N batches of tasks by using the vector operator, including:
[0008] When the second thread bundle group receives the first instruction corresponding to the nth batch of tasks sent by the first thread bundle group and the vector calculation of the (n - 1)th batch of tasks by the second thread bundle group is completed, the second thread bundle group performs vector calculation on the nth batch of tasks by using the vector operator; the first instruction corresponding to the nth batch of tasks is used to identify that the first thread bundle group has completed matrix calculation on the nth batch of tasks by using the first matrix operator.
[0009] Optionally, the second thread bundle group performs vector calculation on the nth batch of tasks by using the vector operator, including:
[0010] The second thread bundle group obtains the first intermediate result corresponding to the nth batch of tasks; the first intermediate result corresponding to the nth batch of tasks is the result of matrix calculation on the nth batch of tasks by the first thread bundle group by using the first matrix operator;
[0011] The second thread bundle group performs vector calculation on the first intermediate result corresponding to the nth batch of tasks by using the vector operator to obtain the second intermediate result corresponding to the nth batch of tasks.
[0012] Optionally, the first thread bundle group performs matrix calculation on the (n + 1)th batch of tasks among the N batches of tasks by using the first matrix operator, including:
[0013] When the first thread bundle group has completed matrix calculation on the nth batch of tasks by using the first matrix operator, the first thread bundle group performs matrix calculation on the (n + 1)th batch of tasks by using the first matrix operator.
[0014] Optionally, the third thread bundle group performs matrix calculation on the (n - 1)th batch of tasks among the N batches of tasks by using the second matrix operator, including:
[0015] When the third thread bundle group receives the second instruction corresponding to the (n - 1)th batch of tasks sent by the second thread bundle group and the third thread bundle group has completed matrix calculation on the (n - 2)th batch of tasks by using the second matrix operator, the third thread bundle group performs matrix calculation on the (n - 1)th batch of tasks by using the second matrix operator; the second instruction corresponding to the (n - 1)th batch of tasks is used to identify that the second thread bundle group has completed vector calculation on the (n - 1)th batch of tasks by using the vector operator.
[0016] Optionally, the third thread bundle group performs matrix calculation on the (n - 1)th batch of tasks by using the second matrix operator, including:
[0017] The third thread bundle group obtains the second intermediate result corresponding to the (n - 1)-th batch of tasks; the second intermediate result corresponding to the (n - 1)-th batch of tasks is the result of the second thread bundle group performing vector calculation on the (n - 1)-th batch of tasks using the vector operator;
[0018] The third thread bundle group performs matrix calculation on the second intermediate result corresponding to the (n - 1)-th batch of tasks using the second matrix operator to obtain the result of the (n - 1)-th batch of tasks.
[0019] Optionally, before the second thread bundle group performs vector calculation on the n-th batch of tasks among the N batches of tasks using the vector operator, it further includes:
[0020] The first thread bundle group finishes matrix calculation on the n-th batch of tasks using the first matrix operator, and the first thread bundle group sends a first instruction corresponding to the n-th batch of tasks to the second thread bundle group.
[0021] On the one hand, an operator execution device provided by an embodiment of the present application includes:
[0022] A matrix calculation module, configured to, when the second thread bundle group performs vector calculation on the n-th batch of tasks among N batches of tasks using the vector operator, perform matrix calculation on the (n + 1)-th batch of tasks among the N batches of tasks by the first thread bundle group using the first matrix operator, and perform matrix calculation on the (n - 1)-th batch of tasks among the N batches of tasks by the third thread bundle group using the second matrix operator; the second thread bundle group includes W vector calculation cores, the first thread bundle group includes S1 tensor calculation cores among S tensor calculation cores, and the third thread bundle group includes S2 tensor calculation cores among the S tensor calculation cores; the value range of n is [2, N], N is an integer greater than 2, both S1 and S2 are less than S, and the sum of S1 and S2 is less than or equal to S.
[0023] Optionally, a vector calculation module, configured to, when the second thread bundle group receives the first instruction corresponding to the n-th batch of tasks sent by the first thread bundle group and the second thread bundle group finishes vector calculation on the (n - 1)-th batch of tasks, perform vector calculation on the n-th batch of tasks by the second thread bundle group using the vector operator; the first instruction corresponding to the n-th batch of tasks is used to indicate that the first thread bundle group finishes matrix calculation on the n-th batch of tasks using the first matrix operator.
[0024] Optionally, the vector calculation module is specifically configured to:
[0025] Obtain the first intermediate result corresponding to the nth batch of tasks through the second warp group; the first intermediate result corresponding to the nth batch of tasks is the result of the first warp group performing matrix calculation on the nth batch of tasks using the first matrix operator;
[0026] Use the vector operator through the second warp group to perform vector calculation on the first intermediate result corresponding to the nth batch of tasks, and obtain the second intermediate result corresponding to the nth batch of tasks.
[0027] Optionally, the matrix calculation module is specifically configured to:
[0028] When the matrix calculation of the nth batch of tasks by the first warp group using the first matrix operator ends, perform matrix calculation on the (n + 1)th batch of tasks by the first warp group using the first matrix operator.
[0029] Optionally, the matrix calculation module is specifically configured to:
[0030] When the third warp group receives the second instruction corresponding to the (n - 1)th batch of tasks sent by the second warp group and the matrix calculation of the (n - 2)th batch of tasks by the third warp group using the second matrix operator ends, perform matrix calculation on the (n - 1)th batch of tasks by the third warp group using the second matrix operator; the second instruction corresponding to the (n - 1)th batch of tasks is used to indicate that the vector calculation of the (n - 1)th batch of tasks by the second warp group using the vector operator ends.
[0031] Optionally, the matrix calculation module is specifically configured to:
[0032] Obtain the second intermediate result corresponding to the (n - 1)th batch of tasks through the third warp group; the second intermediate result corresponding to the (n - 1)th batch of tasks is the result of the second warp group performing vector calculation on the (n - 1)th batch of tasks using the vector operator;
[0033] Use the second matrix operator through the third warp group to perform matrix calculation on the second intermediate result corresponding to the (n - 1)th batch of tasks, and obtain the result of the (n - 1)th batch of tasks.
[0034] Optionally, the matrix calculation module is further configured to: before the second warp group performs vector calculation on the nth batch of tasks among the N batches of tasks using the vector operator, the matrix calculation of the nth batch of tasks by the first warp group using the first matrix operator ends, and the first warp group sends the first instruction corresponding to the nth batch of tasks to the second warp group.
[0035] On the one hand, an embodiment of the present application provides a computer device, including a memory, an artificial intelligence chip, and a computer program stored on the memory and executable on the artificial intelligence chip. When the artificial intelligence chip executes the computer program, the steps of the above-mentioned operator execution method are implemented.
[0036] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned operator execution method.
[0037] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned operator execution method.
[0038] In the embodiment of the present application, during the processing of N batch tasks by the Attention operator, when the vector operator is used to perform vector calculation on the nth batch task in the second thread bundle group, the first thread bundle group uses the first matrix operator to perform matrix calculation on the (n + 1)th batch task, and at the same time, the third thread bundle group uses the second matrix operator to perform matrix calculation on the (n - 1)th batch task. Since both the first thread bundle group and the third thread bundle group are composed of tensor calculation units, the tensor calculation units can process two batch tasks simultaneously. In the above process, the tensor calculation units and the vector calculation units can process different batch tasks simultaneously, avoiding the waiting time for each other, and both the tensor calculation units and the vector calculation units are fully utilized, thereby reducing the time overhead caused by the Attention operator processing multiple batch tasks. In addition, since each thread bundle group only executes one type of calculation, such as the first thread bundle group executes matrix calculation, the second thread bundle group executes vector calculation, and the third thread bundle group executes matrix calculation, the complexity of managing the thread bundle groups can be effectively reduced. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1A It is a schematic structural diagram of an artificial intelligence chip provided by an embodiment of the present application;
[0041] Figure 1B It is a schematic structural diagram of a calculation unit provided by an embodiment of the present application;
[0042] Figure 2 Schematic structural diagram of an Attention operator provided by an embodiment of the present application;
[0043] Figure 3 Schematic structural diagram of the calculation process of an Attention operator provided by an embodiment of the present application;
[0044] Figure 4 Schematic structural diagram of the calculation process of an Attention operator provided by an embodiment of the present application;
[0045] Figure 5 Schematic distribution diagram of a tensor computing core and a vector computing core provided by an embodiment of the present application;
[0046] Figure 6 Schematic distribution diagram of a tensor computing core and a vector computing core provided by an embodiment of the present application;
[0047] Figure 7 Schematic flow diagram of an operator execution method provided by an embodiment of the present application;
[0048] Figure 8 Schematic structural diagram of the calculation process of an Attention operator provided by an embodiment of the present application;
[0049] Figure 9 Schematic structural diagram of the calculation process of an Attention operator provided by an embodiment of the present application;
[0050] Figure 10 Schematic structural diagram of an operator execution device provided by an embodiment of the present application;
[0051] Figure 11 Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Hereinafter, some terms in the embodiments of the present application will be explained to facilitate the understanding of those skilled in the art.
[0054] A warp group is composed of a group of warps. A warp is a scheduling unit when an artificial intelligence chip executes a program. Threads located in a warp can execute the same instruction with different data resources.
[0055] Reference Figure 1A Figure 1A is a structural diagram of an artificial intelligence chip applicable to an embodiment of the present application. The artificial intelligence chip 100 at least includes: a plurality of computing units (Compute Unit, abbreviated as CU) 101, an on-chip cache 105, and a video memory 106. Among them, each computing unit 101 includes: a tensor computing unit 102, a vector computing unit 103, and a register 104. Among them, the tensor computing unit 102 and the vector computing unit 103 are the cores of the computing unit 101.
[0056] The register 104 is a register shared by the tensor computing unit 102 and the vector computing unit 103. The tensor computing unit 102 and the vector computing unit 103 can communicate through the register 104. Compared with the on-chip cache 105, the capacity of the register 104 is smaller than that of the on-chip cache 105, but the data exchange speed is faster than that of the on-chip cache 105.
[0057] The on-chip cache 105 is a temporary memory. Its capacity is smaller than that of the video memory 106, but the data exchange speed is faster than that of the video memory 106.
[0058] The video memory 106 can be a High Bandwidth Memory (abbreviated as HBM), or other types of memories.
[0059] If the capacity of the register 104 can meet the data transmission between the tensor computing unit 102 and the vector computing unit 103, the data between the tensor computing unit 102 and the vector computing unit 103 is transmitted through the register 104.
[0060] If the capacity of the register 104 cannot meet the data transmission between the tensor computing unit 102 and the vector computing unit 103, the data between the tensor computing unit 102 and the vector computing unit 103 is transmitted through the on-chip cache 105.
[0061] If the capacity of the on-chip cache 105 cannot meet the data transmission between the tensor computing unit 102 and the vector computing unit 103, the data between the tensor computing unit 102 and the vector computing unit 103 is transmitted through the video memory 106.
[0062] In addition to including the above structure, the artificial intelligence chip 100 in the present application may further include other structures. In this regard, the present application does not make specific limitations.
[0063] The artificial intelligence chip 100 can be: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a General-purpose Computing on Graphics Processing Units (GPGPU), a Domain Specific Architecture (DSA), etc.
[0064] For Figure 1A the computing unit 101 in Figure 1B an exemplary structure of a computing unit 101 is given. The tensor computing unit 102 in the computing unit 101 includes 4 tensor computing cores (T-Core), and all 4 tensor computing cores are used for tensor computing. The vector computing unit 103 in the computing unit 101 includes 16 vector computing cores (V-Core), and all 16 vector computing cores are used for vector computing.
[0065] In practical applications, one structure of the Attention operator is as Figure 2 shown. The Attention operator includes a first matrix operator, a Scale operator, a Mask operator, a SoftMax operator, and a second matrix operator. Among them, both the first matrix operator and the second matrix operator are matrix operators (MatMul), and the matrix operator can include at least one Matrix Multiply Accumulate (MMA). Among them, the Scale operator, the Mask operator, and the normalization operator are all vector operators.
[0066] In the embodiments of the present application, the processes of training or reasoning any batch of tasks using the Attention operator are the same. For the sake of explanation, hereinafter, taking the use of the Attention operator to reason about any batch of tasks as an example, a detailed explanation will be given.
[0067] When using the Attention operator to reason about any batch of tasks, the specific calculation process of the Attention operator can refer to Figure 3The steps are as follows: First, the tensor calculation unit performs matrix calculations on the query vector Q and the key vector K using the first matrix operator to obtain the correlation between the query vector Q and the key vector K. Then, the vector calculation unit performs vector calculations on the correlation between the query vector Q and the key vector K using vector operators (a scaling operator, a masking operator, and a normalization operator in sequence) to obtain weights. Finally, the tensor calculation unit performs matrix calculations on the weights and the value vector V using the second matrix operator to obtain the final output vector.
[0068] If there are multiple batch tasks to be inferred, the Attention operator can execute the Figure 3 steps shown. To improve the processing efficiency of multiple batch tasks, generally two warp groups are currently used to process each batch task in parallel. Among them, each warp group is responsible for the complete processing process of a batch task. Therefore, each warp group is assigned a tensor calculation unit and a vector calculation unit, and each warp group performs matrix calculations and vector calculations.
[0069] Taking the example of two warp groups processing four batch tasks to be inferred, the specific calculation process can refer to Figure 4 the steps shown. From Figure 4 it can be seen that warp group 1 processes batch task 1 and batch task 3, and warp group 2 processes batch task 2 and batch task 4.
[0070] Taking warp group 1 processing batch task 1 and warp group 2 processing batch task 2 as an example, the tensor calculation unit first performs matrix calculations on batch task 1 using the first matrix operator. Then, the vector calculation unit performs vector calculations on batch task 1 using vector operators. At the same time, the tensor calculation unit performs matrix calculations on batch task 2 using the first matrix operator. Then, the tensor calculation unit performs matrix calculations on batch task 1 using the second matrix operator. At the same time, the vector calculation unit performs vector calculations on batch task 2 using vector operators.
[0071] From the above, it can be seen that although the two warp groups can process each batch task in parallel, when warp group 1 processes a batch task through the tensor calculation unit, warp group 2 cannot simultaneously process a batch task through the tensor calculation unit, that is, the tensor calculation unit cannot process two batch tasks at the same time. This will cause the tensor calculation unit or the vector calculation unit to be idle. Since the tensor calculation unit and the vector calculation unit are not fully utilized, the time overhead caused by the Attention operator inferring the above four batch tasks is relatively large. In addition, each warp group is assigned a tensor calculation unit and a vector calculation unit, and each warp group needs to perform matrix calculations and vector calculations, which will also increase the complexity of managing the warp groups.
[0072] In view of this, when the Attention operator processes N batch tasks, and the Attention operator includes at least a first matrix operator, a vector operator, and a second matrix operator that sequentially execute any batch task, where N is an integer greater than 2, in order to improve the processing speed of the Attention operator for N batch tasks and reduce the time overhead, this application proposes an operator execution method.
[0073] Combined with the Figure 1A system architecture diagram shown in this application, it is set that Figure 1A the tensor calculation unit in the artificial intelligence chip in includes S tensor calculation cores, and the vector calculation unit includes W vector calculation cores, where S and W are both positive integers. Select S1 tensor calculation cores from the S tensor calculation cores and allocate them to the first thread bundle group, and select S2 tensor calculation cores from the S tensor calculation cores and allocate them to the third thread bundle group. Among them, the first thread bundle group is used to perform matrix calculations on batch tasks using the first matrix operator, and the third thread bundle group is used to perform matrix calculations on batch tasks using the second matrix operator. There is no overlap between the S1 tensor calculation cores and the S2 tensor calculation cores, both S1 and S2 are less than S, and the sum of S1 and S2 is less than or equal to S.
[0074] In addition, allocate the W vector calculation cores in the artificial intelligence chip to the second thread bundle group, where the second thread bundle group is used to perform vector calculations on batch tasks using the vector operator.
[0075] For example, it is set that the artificial intelligence chip includes 4 computing units, and each computing unit includes 4 tensor calculation cores and 16 vector calculation cores. The allocation scheme of the tensor calculation cores and vector calculation cores in the above artificial intelligence chip can refer to Figure 5 , that is, allocate the tensor calculation cores in computing units 1 and 2 to the first thread bundle group, allocate the tensor calculation cores in computing units 3 and 4 to the third thread bundle group, and allocate the vector calculation cores in computing units 1, 2, 3, and 4 to the second thread bundle group.
[0076] The allocation scheme of the tensor calculation cores and vector calculation cores in the above artificial intelligence chip can also refer to Figure 6 , that is, evenly allocate the 4 tensor calculation cores in each computing unit to the first thread bundle group and the third thread bundle group, and allocate the vector calculation cores in computing units 1, 2, 3, and 4 to the second thread bundle group.
[0077] The allocation scheme of the tensor calculation cores and vector calculation cores in the above artificial intelligence chip can also adopt other allocation schemes, which will not be elaborated here. Since the tensor calculation cores in the same computing unit are allocated to the same thread bundle group, the processing efficiency of the thread bundle group can be higher. Therefore, in the following, all are in accordance withFigure 5 The shown allocation scheme allocates tensor calculation cores and vector calculation cores in the artificial intelligence chip.
[0078] The process of an operator execution method proposed in an embodiment of this application is executed interactively by the above-mentioned first warp group, second warp group, and third warp group, and the tensor calculation cores or vector calculation cores allocated to the first warp group, second warp group, and third warp group refer to Figure 5 , and the operator execution method may include the following steps:
[0079] When the second warp group performs vector calculation on the nth batch task among N batch tasks using a vector operator, the first warp group performs matrix calculation on the (n + 1)th batch task among N batch tasks using a first matrix operator, and the third warp group performs matrix calculation on the (n - 1)th batch task among N batch tasks using a second matrix operator; the value range of n is [2, N].
[0080] Specifically, since the second warp group performs vector calculation on the nth batch task using a vector operator and needs to rely on the first warp group performing matrix calculation on the nth batch task using a first matrix operator, before the second warp group performs vector calculation on the nth batch task, the nth batch task has already been subjected to matrix calculation by the first matrix operator and the calculation result has been obtained, and the (n - 1)th batch task has already been subjected to vector calculation by the vector operator and the calculation result has been obtained. Therefore, when the second warp group performs vector calculation on the nth batch task using a vector operator, the first warp group at this time can perform matrix calculation on the (n + 1)th batch task using a first matrix operator, and the third warp group at this time can perform matrix calculation on the (n - 1)th batch task using a second matrix operator.
[0081] It should be understood that in the embodiment of this application, the nth batch task may be a task to be processed by an Attention operator, may be a task to be processed by a multi-head attention mechanism (Multi-Head Attention, abbreviated as MHA) operator, may also be a task to be processed by an Attention operator in the MHA operator, or may also be a partial task to be processed by an Attention operator in the MHA operator, which is not limited herein.
[0082] For the convenience of subsequent description, the following will take the nth batch task as a task to be processed by an Attention operator as an example for detailed description.
[0083] In the above method, during the processing of N batches of tasks by the Attention operator, when the vector operator is used to perform vector calculation on the n-th batch of tasks in the second warp group, the first warp group uses the first matrix operator to perform matrix calculation on the (n + 1)-th batch of tasks, and at the same time, the third warp group uses the second matrix operator to perform matrix calculation on the (n - 1)-th batch of tasks. Since both the first warp group and the third warp group are composed of tensor calculation units, the tensor calculation units can process two batches of tasks simultaneously. In the above process, the tensor calculation units and the vector calculation units can process different batches of tasks simultaneously, avoiding the waiting time for each other, and both the tensor calculation units and the vector calculation units are fully utilized, thereby reducing the time overhead caused by the Attention operator processing multiple batches of tasks. In addition, since each warp group only executes one type of calculation, such as the first warp group executes matrix calculation, the second warp group executes vector calculation, and the third warp group executes matrix calculation, the complexity of managing warp groups can be effectively reduced.
[0084] For each batch of tasks, the execution processes of the first warp group, the second warp group, and the third warp group are introduced in detail below, specifically including the following embodiments as Figure 7 shown.
[0085] S701, when the first warp group is in an idle state, the first warp group uses the first matrix operator to perform matrix calculation on the n-th batch of tasks.
[0086] S702, after the first warp group finishes performing matrix calculation on the n-th batch of tasks using the first matrix operator, the first warp group sends the first instruction corresponding to the n-th batch of tasks to the second warp group, where the first instruction corresponding to the n-th batch of tasks is used to indicate that the first warp group has finished performing matrix calculation on the n-th batch of tasks using the first matrix operator.
[0087] In the above S702, after the first warp group finishes performing matrix calculation on the n-th batch of tasks using the first matrix operator, the first intermediate result corresponding to the n-th batch of tasks can be obtained and written into the register.
[0088] S703, if the second warp group receives the first instruction corresponding to the n-th batch of tasks and the vector calculation of the (n - 1)-th batch of tasks by the second warp group is completed, the second warp group uses the vector operator to perform vector calculation on the n-th batch of tasks.
[0089] S704. After the second thread bundle group finishes vector calculation on the nth batch of tasks using a vector operator, the second thread bundle group sends a second instruction corresponding to the nth batch of tasks to the third thread bundle group. The second instruction corresponding to the nth batch of tasks is used to indicate that the second thread bundle group has finished vector calculation on the nth batch of tasks using a vector operator.
[0090] In S704 above, the second thread bundle group first obtains the first intermediate result corresponding to the nth batch of tasks. The first intermediate result corresponding to the nth batch of tasks is the result of the first thread bundle group performing matrix calculation on the nth batch of tasks using a first matrix operator. The second thread bundle group then performs vector calculation on the first intermediate result corresponding to the nth batch of tasks using a vector operator to obtain a second intermediate result corresponding to the nth batch of tasks, and writes the second intermediate result corresponding to the nth batch of tasks into a register.
[0091] In S704 above, the second thread bundle group can obtain the first intermediate result corresponding to the nth batch of tasks through implementation method A1, A2 or A3.
[0092] Implementation method A1. If the register into which the first intermediate result corresponding to the nth batch of tasks is written is in the same computing unit as the second thread bundle group, the second thread bundle group can obtain the first intermediate result corresponding to the nth batch of tasks from the register.
[0093] Implementation method A2. If the register into which the first intermediate result corresponding to the nth batch of tasks is written is in a different computing unit from the second thread bundle group, and the second thread bundle group can access the register in other computing units, the second thread bundle group can obtain the first intermediate result corresponding to the nth batch of tasks from the register.
[0094] Implementation method A3. If the register into which the first intermediate result corresponding to the nth batch of tasks is written is in a different computing unit from the second thread bundle group, and the second thread bundle group cannot access the register in other computing units, the first thread bundle group first writes the first intermediate result corresponding to the nth batch of tasks into the on-chip cache, and then the second thread bundle group obtains the first intermediate result corresponding to the nth batch of tasks from the on-chip cache.
[0095] S705. When the second thread bundle group performs vector calculation on the nth batch of tasks using a vector operator, the first thread bundle group performs matrix calculation on the (n + 1)th batch of tasks using a first matrix operator.
[0096] In S705 above, when the second thread bundle group performs vector calculation on the nth batch of tasks using a vector operator, the matrix calculation of the nth batch of tasks by the first thread bundle group has ended. Therefore, the first thread bundle group can perform matrix calculation on the (n + 1)th batch of tasks using a first matrix operator.
[0097] S706, when performing vector calculation on the nth batch of tasks using a vector operator in the second thread bundle group, if the third thread bundle group receives the second instruction corresponding to the (n - 1)th batch of tasks sent by the second thread bundle group and the matrix calculation of the (n - 2)th batch of tasks by the third thread bundle group using the second matrix operator is completed, then the third thread bundle group performs matrix calculation on the (n - 1)th batch of tasks using the second matrix operator.
[0098] In the above 706, the third thread bundle group obtains the second intermediate result corresponding to the (n - 1)th batch of tasks from the register, where the second intermediate result corresponding to the (n - 1)th batch of tasks is the result of the second thread bundle group performing vector calculation on the (n - 1)th batch of tasks using the vector operator; the third thread bundle group performs matrix calculation on the second intermediate result corresponding to the (n - 1)th batch of tasks using the second matrix operator to obtain the result of the (n - 1)th batch of tasks, and writes the result of the (n - 1)th batch of tasks into the register. To facilitate the output of the result of the (n - 1)th batch of tasks, the third thread bundle group can also write the result of the (n - 1)th batch of tasks into the on-chip cache.
[0099] It should be understood that when the nth batch of tasks is the task to be processed by the Attention operator or the task to be processed by the MHA operator, the result of the nth batch of tasks in the on-chip cache can be directly used as the final result of the nth batch of tasks. When the nth batch of tasks is the task to be processed by an Attention operator in the MHA operator or a partial task to be processed by an Attention operator in the MHA operator, the results of each nth batch of tasks in the on-chip cache can be concatenated to obtain the final result of the nth batch of tasks.
[0100] The above Figure 7 The steps shown provide a method for the first thread bundle group, the second thread bundle group, and the third thread bundle group to process each batch of tasks in parallel.
[0101] For ease of understanding, the following takes the Attention operator reasoning on 3 batches of tasks as an example to specifically discuss the above operator execution method in this application. Since the calculation durations of the first matrix operator, the second matrix operator, and the vector operator are not the same, for ease of discussion, it is assumed in this application that the calculation durations of the first matrix operator and the second matrix operator are the same. The following specifically discusses the above operator execution method in this application in combination with the relationships between the calculation durations corresponding to the first matrix operator, the second matrix operator, and the vector operator, including the following implementation manners B1 or B2.
[0102] Implementation manner B1, the calculation duration of the vector operator is less than or equal to the calculation duration of the first matrix operator (or the second matrix operator). The Attention operator calculates for 3 batches of tasks in sequence, and the specific calculation process can refer toFigure 8 The steps shown.
[0103] Assume that the first warp group, the second warp group, and the third warp group are all in the idle state at time t0, and the instruction transmission between the first warp group, the second warp group, and the third warp group does not take time.
[0104] At time t0, the first warp group performs matrix calculations on the query vector Q and the key vector K of the first batch of tasks using the first matrix operator, obtains the correlation between the query vector Q and the key vector K of the first batch of tasks at time t1, and sends the first instruction corresponding to the first batch of tasks to the second warp group at time t1. Among them, the first instruction corresponding to the first batch of tasks is used to indicate that the first warp group has completed the matrix calculation of the first batch of tasks using the first matrix operator.
[0105] At time t1, the second warp group receives the first instruction corresponding to the first batch of tasks, performs vector calculations on the correlation corresponding to the first batch of tasks using the vector operator at time t1, obtains the weight corresponding to the first batch of tasks at time t2, and sends the second instruction corresponding to the first batch of tasks to the third warp group at time t2. Among them, the second instruction corresponding to the first batch of tasks is used to indicate that the second warp group has completed the vector calculation of the first batch of tasks using the vector operator.
[0106] When the second warp group performs vector calculations on the correlation corresponding to the first batch of tasks using the vector operator, at time t1, the first warp group performs matrix calculations on the query vector Q and the key vector K of the second batch of tasks using the first matrix operator, obtains the correlation between the query vector Q and the key vector K of the second batch of tasks at time t3, and sends the first instruction corresponding to the second batch of tasks to the second warp group at time t3. Among them, the first instruction corresponding to the second batch of tasks is used to indicate that the first warp group has completed the matrix calculation of the second batch of tasks using the first matrix operator.
[0107] At time t3, the second warp group receives the first instruction corresponding to the second batch of tasks and the vector calculation of the first batch of tasks ends at time t2. Since time t3 is greater than time t2, therefore, at time t3, the second warp group performs vector calculations on the correlation corresponding to the second batch of tasks using the vector operator, obtains the weight corresponding to the second batch of tasks at time t4, and sends the second instruction corresponding to the second batch of tasks to the third warp group at time t4. Among them, the second instruction corresponding to the second batch of tasks is used to indicate that the second warp group has completed the vector calculation of the second batch of tasks using the vector operator.
[0108] When the second thread bundle group performs vector calculation on the relevance corresponding to the second batch of tasks using a vector operator, the third thread bundle group performs matrix calculation on the weights corresponding to the first batch of tasks and the value vector V of the first batch of tasks using a second matrix operator at time t2, and obtains the output vector of the first batch of tasks at time t4. Moreover, the first thread bundle group performs matrix calculation on the query vector Q and the key vector K of the third batch of tasks using a first matrix operator at time t3, obtains the relevance between the query vector Q and the key vector K of the third batch of tasks at time t5, and sends a first instruction corresponding to the third batch of tasks to the second thread bundle group at time t5, where the first instruction corresponding to the third batch of tasks is used to indicate that the first thread bundle group has completed the matrix calculation of the third batch of tasks using the first matrix operator.
[0109] The second thread bundle group receives the first instruction corresponding to the third batch of tasks at time t5 and finishes the vector calculation of the second batch of tasks at time t4. Since time t5 is greater than time t4, the second thread bundle group performs vector calculation on the relevance corresponding to the third batch of tasks using a vector operator at time t5, obtains the weights corresponding to the third batch of tasks at time t6, and sends a second instruction corresponding to the third batch of tasks to the third thread bundle group at time t6, where the second instruction corresponding to the third batch of tasks is used to indicate that the second thread bundle group has completed the vector calculation of the third batch of tasks using the vector operator.
[0110] When the second thread bundle group performs vector calculation on the relevance corresponding to the third batch of tasks using a vector operator, the third thread bundle group performs matrix calculation on the weights corresponding to the second batch of tasks and the value vector V of the second batch of tasks using a second matrix operator at time t4, and obtains the output vector of the second batch of tasks at time t6.
[0111] The third thread bundle group receives the second instruction corresponding to the third batch of tasks at time t6 and finishes the matrix calculation of the second batch of tasks at time t6. Therefore, the third thread bundle group performs matrix calculation on the weights corresponding to the third batch of tasks and the value vector V of the third batch of tasks using a second matrix operator at time t6, and obtains the output vector of the third batch of tasks at time t8.
[0112] In Embodiment B2, the calculation duration of the vector operator is greater than that of the first matrix operator (or the second matrix operator). The Attention operator calculates for 3 batches of tasks in sequence, and the specific calculation process can refer to Figure 9 the steps shown.
[0113] Assume that the first thread bundle group, the second thread bundle group, and the third thread bundle group are all in an idle state at time t0, and the instruction transmission between the first thread bundle group, the second thread bundle group, and the third thread bundle group does not take time.
[0114] At time t0, the first thread warp group performs matrix calculations on the query vector Q and the key vector K of the first batch of tasks using the first matrix operator, obtains the correlation between the query vector Q and the key vector K of the first batch of tasks at time t1, and sends the first instruction corresponding to the first batch of tasks to the second thread warp group at time t1. The first instruction corresponding to the first batch of tasks is used to indicate that the first thread warp group has completed the matrix calculation of the first batch of tasks using the first matrix operator.
[0115] At time t1, the second thread warp group receives the first instruction corresponding to the first batch of tasks, performs vector calculations on the correlation corresponding to the first batch of tasks using the vector operator at time t1, obtains the weight corresponding to the first batch of tasks at time t3, and sends the second instruction corresponding to the first batch of tasks to the third thread warp group at time t3. The second instruction corresponding to the first batch of tasks is used to indicate that the second thread warp group has completed the vector calculation of the first batch of tasks using the vector operator.
[0116] When the second thread warp group performs vector calculations on the correlation corresponding to the first batch of tasks using the vector operator, the first thread warp group performs matrix calculations on the query vector Q and the key vector K of the second batch of tasks using the first matrix operator at time t1, obtains the correlation between the query vector Q and the key vector K of the second batch of tasks at time t2, and sends the first instruction corresponding to the second batch of tasks to the second thread warp group at time t2. The first instruction corresponding to the second batch of tasks is used to indicate that the first thread warp group has completed the matrix calculation of the second batch of tasks using the first matrix operator.
[0117] At time t2, the second thread warp group receives the first instruction corresponding to the second batch of tasks and finishes the vector calculation of the first batch of tasks at time t3. Since time t3 is greater than time t2, the second thread warp group performs vector calculations on the correlation corresponding to the second batch of tasks using the vector operator at time t3, obtains the weight corresponding to the second batch of tasks at time t6, and sends the second instruction corresponding to the second batch of tasks to the third thread warp group at time t6. The second instruction corresponding to the second batch of tasks is used to indicate that the second thread warp group has completed the vector calculation of the second batch of tasks using the vector operator.
[0118] When the second thread warp group performs vector calculation on the relevance corresponding to the second batch of tasks using a vector operator, the third thread warp group performs matrix calculation on the weight corresponding to the first batch of tasks and the value vector V of the first batch of tasks using a second matrix operator at time t3, and obtains the output vector of the first batch of tasks at time t5. Moreover, the first thread warp group performs matrix calculation on the query vector Q and the key vector K of the third batch of tasks using a first matrix operator at time t2, obtains the relevance between the query vector Q and the key vector K of the third batch of tasks at time t4, and sends a first instruction corresponding to the third batch of tasks to the second thread warp group at time t4. Here, the first instruction corresponding to the third batch of tasks is used to indicate that the first thread warp group has completed the matrix calculation of the third batch of tasks using the first matrix operator.
[0119] The second thread warp group receives the first instruction corresponding to the third batch of tasks at time t4 and finishes the vector calculation of the second batch of tasks at time t6. Since time t6 is greater than time t4, the second thread warp group performs vector calculation on the relevance corresponding to the third batch of tasks using a vector operator at time t6, obtains the weight corresponding to the third batch of tasks at time t8, and sends a second instruction corresponding to the third batch of tasks to the third thread warp group at time t8. Here, the second instruction corresponding to the third batch of tasks is used to indicate that the second thread warp group has completed the vector calculation of the third batch of tasks using the vector operator.
[0120] When the second thread warp group performs vector calculation on the relevance corresponding to the third batch of tasks using a vector operator, the third thread warp group performs matrix calculation on the weight corresponding to the second batch of tasks and the value vector V of the second batch of tasks using a second matrix operator at time t6, and obtains the output vector of the second batch of tasks at time t7.
[0121] The third thread warp group receives the second instruction corresponding to the third batch of tasks at time t8 and finishes the matrix calculation of the second batch of tasks at time t7. Since time t8 is greater than time t7, the third thread warp group performs matrix calculation on the weight corresponding to the third batch of tasks and the value vector V of the third batch of tasks using a second matrix operator at time t8, and obtains the output vector of the third batch of tasks at time t9.
[0122] In the above-mentioned Embodiment A1 or Embodiment A2, any moment from t0 to t9 in different embodiments can be the same moment or different moments.
[0123] Based on the same technical concept, an embodiment of the present application provides a structural schematic diagram of an operator execution device, as Figure 10 shown. The operator execution device 1000 includes:
[0124] The matrix calculation module 1001 is configured to, when the second thread bundle group performs vector calculation on the nth batch task among N batch tasks using a vector operator, perform matrix calculation on the (n + 1)th batch task among the N batch tasks using a first matrix operator through the first thread bundle group, and perform matrix calculation on the (n - 1)th batch task among the N batch tasks using a second matrix operator through the third thread bundle group; the second thread bundle group includes W vector calculation cores, the first thread bundle group includes S1 tensor calculation cores among S tensor calculation cores, the third thread bundle group includes S2 tensor calculation cores among the S tensor calculation cores, and there is no overlap between the S1 tensor calculation cores and the S2 tensor calculation cores; the value range of n is [2, N], N is an integer greater than 2, both S1 and S2 are less than S, and the sum of S1 and S2 is less than or equal to S.
[0125] Optionally, the vector calculation module 1002 is configured to, when the second thread bundle group receives the first instruction corresponding to the nth batch task sent by the first thread bundle group and the vector calculation of the (n - 1)th batch task by the second thread bundle group ends, perform vector calculation on the nth batch task using the vector operator through the second thread bundle group; the first instruction corresponding to the nth batch task is used to indicate that the first thread bundle group has completed matrix calculation on the nth batch task using the first matrix operator.
[0126] Optionally, the vector calculation module 1002 is specifically configured to:
[0127] Obtain the first intermediate result corresponding to the nth batch task through the second thread bundle group; the first intermediate result corresponding to the nth batch task is the result of the first thread bundle group performing matrix calculation on the nth batch task using the first matrix operator;
[0128] Perform vector calculation on the first intermediate result corresponding to the nth batch task using the vector operator through the second thread bundle group to obtain the second intermediate result corresponding to the nth batch task.
[0129] Optionally, the matrix calculation module 1001 is specifically configured to:
[0130] When the first thread bundle group finishes matrix calculation on the nth batch task using the first matrix operator, perform matrix calculation on the (n + 1)th batch task using the first matrix operator through the first thread bundle group.
[0131] Optionally, the matrix calculation module 1001 is specifically configured to:
[0132] When the third thread bundle group receives the second instruction corresponding to the (n - 1)-th batch of tasks sent by the second thread bundle group and the matrix calculation of the (n - 2)-th batch of tasks by the third thread bundle group using the second matrix operator is completed, the third thread bundle group performs matrix calculation on the (n - 1)-th batch of tasks using the second matrix operator; the second instruction corresponding to the (n - 1)-th batch of tasks is used to indicate that the vector calculation of the (n - 1)-th batch of tasks by the second thread bundle group using the vector operator is completed.
[0133] Optionally, the matrix calculation module 1001 is specifically configured to:
[0134] Obtain, by the third thread bundle group, the second intermediate result corresponding to the (n - 1)-th batch of tasks; the second intermediate result corresponding to the (n - 1)-th batch of tasks is the result of the vector calculation of the (n - 1)-th batch of tasks by the second thread bundle group using the vector operator;
[0135] Perform matrix calculation on the second intermediate result corresponding to the (n - 1)-th batch of tasks by the third thread bundle group using the second matrix operator to obtain the result of the (n - 1)-th batch of tasks.
[0136] Optionally, the matrix calculation module 1001 is further configured to: before the second thread bundle group performs vector calculation on the n-th batch of tasks among the N batches of tasks using the vector operator, complete the matrix calculation of the n-th batch of tasks by the first thread bundle group using the first matrix operator, and send, by the first thread bundle group, the first instruction corresponding to the n-th batch of tasks to the second thread bundle group.
[0137] Based on the same technical concept, an embodiment of the present application provides a computer device, as Figure 11 shown, including at least one artificial intelligence chip 1101 and a memory 1102 connected to the at least one artificial intelligence chip. In the embodiment of the present application, the specific connection medium between the artificial intelligence chip 1101 and the memory 1102 is not limited. Figure 11 Taking the connection between the artificial intelligence chip 1101 and the memory 1102 through a bus as an example. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0138] In the embodiment of the present application, the memory 1102 stores instructions executable by the at least one artificial intelligence chip 1101. The at least one artificial intelligence chip 1101 can execute the steps of the above operator execution method by executing the instructions stored in the memory 1102.
[0139] Among them, the artificial intelligence chip 1101 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and circuits, and implement operator execution by running or executing instructions stored in the memory 1102 and calling data stored in the memory 1102. Optionally, the artificial intelligence chip 1101 may include one or more processing units. The artificial intelligence chip 1101 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the artificial intelligence chip 1101. In some embodiments, the artificial intelligence chip 1101 and the memory 1102 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.
[0140] The artificial intelligence chip 1101 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0141] The memory 1102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1102 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 1102 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 1102 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0142] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium that stores a computer program executable by a computer device. When the program runs on the computer device, it causes the computer device to execute the steps of the above-mentioned operator execution method.
[0143] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, they cause the computer device to execute the steps of the above-mentioned operator execution method.
[0144] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0145] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in a flow or flows in the flowchart and / or a block or blocks in the block diagram.
[0146] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including instruction means that implement the functions specified in a flow or flows in the flowchart and / or a block or blocks in the block diagram.
[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a flow or flows in the flowchart and / or a block or blocks in the block diagram.
[0148] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. An operator execution method, characterized in that: Applicable to an artificial intelligence chip including S tensor computing cores and W vector computing cores, wherein the artificial intelligence chip is used to process N batches of tasks using an attention mechanism operator; The attention mechanism operator at least includes a first matrix operator, a vector operator and a second matrix operator that sequentially execute any batch of tasks in the N batches of tasks; S and W are both positive integers, and N is an integer greater than 2; After the second thread warp group receives the first instruction corresponding to the nth batch of tasks sent by the first thread warp group and the second thread warp group completes vector calculation on the n-1th batch of tasks, the first instruction corresponding to the nth batch of tasks is used to indicate that the first thread warp group has completed matrix calculation on the nth batch of tasks using the first matrix operator, and the second thread warp group obtains the first intermediate result corresponding to the nth batch of tasks; the first intermediate result corresponding to the nth batch of tasks is the result of the first thread warp group performing matrix calculation on the nth batch of tasks using the first matrix operator, and the second thread warp group performs vector calculation on the first intermediate result corresponding to the nth batch of tasks using the vector operator to obtain the nth batch of tasks. When the second intermediate result corresponding to the task is obtained, the first thread warp group uses the first matrix operator to perform matrix calculation on the n+1th batch task among the N batch tasks, and the third thread warp group uses the second matrix operator to perform matrix calculation on the n-1th batch task among the N batch tasks; the second thread warp group includes the W vector computing cores, the first thread warp group includes S1 tensor computing cores among the S tensor computing cores, and the third thread warp group includes S2 tensor computing cores among the S tensor computing cores, and there is no overlap between the S1 tensor computing cores and the S2 tensor computing cores; the value range of n is [2,N], S1 and S2 are both less than S, and the sum of S1 and S2 is less than or equal to S.
2. The method according to claim 1, characterized in that The first warp group uses the first matrix operator to perform matrix calculation on the (n+1)th batch of tasks in the N batches of tasks, including: After the first warp group completes matrix calculation on the nth batch of tasks using the first matrix operator, the first warp group performs matrix calculation on the (n+1)th batch of tasks using the first matrix operator.
3. The method according to claim 1, characterized in that The third warp group uses the second matrix operator to perform matrix calculation on the n-1th batch of tasks in the N batches of tasks, including: When the third thread warp group receives the second instruction corresponding to the n-1th batch of tasks sent by the second thread warp group and the third thread warp group completes matrix calculation on the n-2th batch of tasks using the second matrix operator, the third thread warp group performs matrix calculation on the n-1th batch of tasks using the second matrix operator; the second instruction corresponding to the n-1th batch of tasks is used to indicate that the second thread warp group completes vector calculation on the n-1th batch of tasks using the vector operator.
4. The method according to claim 3, characterized in that The third warp group uses the second matrix operator to perform matrix calculation on the n-1th batch of tasks, including: The third warp group obtains the second intermediate result corresponding to the n-1th batch of tasks; the second intermediate result corresponding to the n-1th batch of tasks is a result of the second warp group performing vector calculation on the n-1th batch of tasks using the vector operator; The third thread warp group uses the second matrix operator to perform matrix calculation on the second intermediate result corresponding to the n-1th batch of tasks to obtain the result of the n-1th batch of tasks.
5. The method according to claim 1, characterized in that Before the second warp group uses the vector operator to perform vector calculation on the nth batch of tasks in the N batches of tasks, the method further includes: The first warp group completes matrix calculation of the nth batch of tasks by using the first matrix operator, and the first warp group sends a first instruction corresponding to the nth batch of tasks to the second warp group.
6. A computer device comprising a memory, the artificial intelligence chip, and a computer program stored in the memory and running on the artificial intelligence chip, characterized in that: When the artificial intelligence chip executes the computer program, the steps of the method as described in any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium, characterized in that: It stores a computer program executable by a computer device, and when the computer program is run on the computer device, the computer device executes the steps of any method as claimed in claims 1-5.
8. A computer program product, characterized in that The computer program product comprises a computer program stored on a computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Operator execution method and device, storage medium and program product
CN118132156A