Parallel operation method of operator flow, computer equipment and readable storage medium
By dividing the operator flow in large-model inference into multiple execution stages and selecting appropriate parallel strategies for each stage, the problem of imbalance in computing power utilization is solved, and more efficient computing speed and lower latency are achieved.
Patent Information
- Application Number
- CN202510653876.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the process of large-scale model inference, a single data parallel or tensor parallel strategy leads to uneven computing power utilization, resulting in idle hardware resources and difficult to meet the low-latency requirements.
The target operator flow is divided into multiple execution stages. Each execution stage adopts an appropriate parallel strategy (regular parallelism, data parallelism, tensor parallelism or mixed parallelism). After completing the operation at each execution stage, the operation result is written directly into the memory resources of the target artificial intelligence chip according to the parallel strategy of the next execution stage.
By accurately adapting parallel strategies and optimizing the direct writing method of operation results, the operator resources are fully utilized, the computing speed is significantly improved, the inference delay of large models is reduced, and the inference efficiency is improved.
Smart Images

Figure CN120179295A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence chips, and particularly to a method for parallel running of operator streams, a computer device, and a readable storage medium. Background Art
[0002] In the scenario of large model inference, efficiently utilizing hardware resources to reduce inference latency is the key to improving model performance. Data parallelism and tensor parallelism are two common strategies for accelerating large model calculations. Among them, data parallelism divides input data onto different computing devices (such as GPUs (Graphics Processing Units)) for parallel processing, and tensor parallelism splits model weights across computing devices for collaborative computing. Both aim to fully utilize the computing power of operators.
[0003] However, the large model inference process involves multiple computing links, and there may be significant differences in the matrix calculation scales and characteristics involved in each link. In response to this situation, if a single data parallelism or tensor parallelism strategy is adopted, it often leads to unbalanced utilization of computing power. For example, for the computing tasks corresponding to certain operators, hardware resources can be fully utilized; but for the computing tasks corresponding to other operators, due to the data partitioning or tensor splitting method not matching the computing tasks of the operator, a large amount of computing power will be idle, making it impossible to achieve efficient utilization of computing resources and difficult to meet the strict requirements of large model inference for low latency. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method for parallel running of operator streams, a computer device, and a readable storage medium that can reduce large model inference latency.
[0005] In a first aspect, this application provides a method for parallel running of operator streams, including:
[0006] Dividing a target operator stream into multiple execution stages, each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy;
[0007] During the running process of the target operator stream, for any one of the execution stages, adopt the parallel strategy corresponding to the execution stage inside each artificial intelligence chip to execute the operations corresponding to the operators in the execution stage, and after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, directly write part or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip, and the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips;
[0008] After the calculations are completed in each execution stage, the operation result of the target operator stream is obtained based on the final operation results of each of the artificial intelligence chips.
[0009] In one embodiment, the conventional parallel strategy includes each of the artificial intelligence chips performing the operation corresponding to the operator on the full data matrix and the full weight matrix;
[0010] The hybrid parallel strategy includes: when the data matrix is the left matrix in a multiplication operation, splitting the data matrix into multiple data sub - matrices with columns as the splitting dimension, and splitting the weight matrix into multiple weight sub - matrices with rows as the splitting dimension; or, when the data matrix is the right matrix in a multiplication operation, splitting the data matrix into multiple data sub - matrices with rows as the splitting dimension, and splitting the weight matrix into multiple weight sub - matrices with columns as the splitting dimension; each artificial intelligence chip performs the operation corresponding to the operator based on the assigned data sub - matrix and weight sub - matrix.
[0011] In one embodiment, the execution stage includes at least one of the following:
[0012] The first execution stage, the first execution stage includes at least one first operator, the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirement, and the first execution stage corresponds to the conventional parallel strategy;
[0013] The second execution stage, the second execution stage includes at least one second operator, the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirement, but the weight matrix of the operation satisfies the parallel requirement, and the second execution stage corresponds to the tensor parallel strategy;
[0014] The third execution stage, the third execution stage includes at least one third operator, the third operator is an operator for which the data matrix of the operation satisfies the parallel requirement, but the weight matrix of the operation does not satisfy the parallel requirement, and the third execution stage corresponds to the data parallel strategy;
[0015] The fourth execution stage, the fourth execution stage includes at least one fourth operator, the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirement, and the fourth execution stage corresponds to the hybrid parallel strategy.
[0016] In one embodiment, when the parallel strategy corresponding to the execution stage is the conventional parallel strategy, based on the parallel strategy corresponding to the next execution stage, writing part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0017] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy, for any one of the artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is retained in the memory resource of the artificial intelligence chip itself;
[0018] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with the row as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of first data sub-matrices corresponding to each artificial intelligence chip respectively, and the first data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resource of the artificial intelligence chip itself;
[0019] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively, and the second data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resource of the artificial intelligence chip itself, wherein the division method of the second data sub-matrix includes:
[0020] When the data matrix in the next execution stage is the left matrix of the multiplication operation, with the column as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, with the row as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively.
[0021] In one embodiment, the parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0022] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is written into the memory resources of each artificial intelligence chip respectively;
[0023] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, according to the number of artificial intelligence chips, with columns as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple third data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves.
[0024] In one embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0025] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0026] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fourth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0027] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fifth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0028] In one embodiment, the parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes:
[0029] Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction result:
[0030] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, for any one of the artificial intelligence chips, retain the full reduction result in the memory resources of the artificial intelligence chip itself;
[0031] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, divide the full reduction result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip in terms of the row division dimension according to the number of artificial intelligence chips, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself.
[0032] In a second aspect, the present application further provides a parallel operation device for an operator stream, including:
[0033] A division module, configured to divide a target operator stream into multiple execution stages, each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy;
[0034] An execution module, configured to, during the operation of the target operator stream, for any one of the execution stages, adopt the parallel strategy corresponding to the execution stage inside each artificial intelligence chip to execute the operations corresponding to the operators in the execution stage, and after the execution stage operation is completed, based on the parallel strategy corresponding to the next execution stage, write part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip, where the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips;
[0035] A determination module, configured to obtain the operation result of the target operator stream based on the final operation results of each artificial intelligence chip after all execution stages have completed calculations.
[0036] In one embodiment, the conventional parallel strategy includes each of the artificial intelligence chips performing the operation corresponding to the operator on the full data matrix and the full weight matrix;
[0037] The hybrid parallel strategy: when the data matrix is the left matrix of the multiplication operation, the data matrix is split into multiple data sub - matrices with columns as the splitting dimension, and the weight matrix is split into multiple weight sub - matrices with rows as the splitting dimension; or, when the data matrix is the right matrix of the multiplication operation, the data matrix is split into multiple data sub - matrices with rows as the splitting dimension, and the weights are split into multiple weight sub - matrices with columns as the splitting dimension; each artificial intelligence chip performs the operation corresponding to the operator based on the allocated data sub - matrix and the weight sub - matrix.
[0038] In one embodiment, the execution phase includes at least one of the following:
[0039] The first execution phase, the first execution phase includes at least one first operator, the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirement, and the first execution phase corresponds to the conventional parallel strategy;
[0040] The second execution phase, the second execution phase includes at least one second operator, the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirement, but the weight matrix of the operation satisfies the parallel requirement, and the second execution phase corresponds to the tensor parallel strategy;
[0041] The third execution phase, the third execution phase includes at least one third operator, the third operator is an operator for which the data matrix of the operation satisfies the parallel requirement, but the weight matrix of the operation does not satisfy the parallel requirement, and the third execution phase corresponds to the data parallel strategy;
[0042] The fourth execution phase, the fourth execution phase includes at least one fourth operator, the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirement, and the fourth execution phase corresponds to the hybrid parallel strategy.
[0043] In one embodiment, the parallel strategy corresponding to the execution phase is the conventional parallel strategy. Based on the parallel strategy corresponding to the next execution phase, writing part or all of the operation results of each artificial intelligence chip in the execution phase directly into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0044] If the parallel strategy corresponding to the next execution phase is the tensor parallel strategy, then for any artificial intelligence chip, the operation result of the artificial intelligence chip in the execution phase is retained in the memory resource of the artificial intelligence chip itself;
[0045] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with the row division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple first data sub-matrices respectively corresponding to each artificial intelligence chip, and retain the first data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself;
[0046] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices respectively corresponding to each artificial intelligence chip, and retain the second data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself, wherein the division method of the second data sub-matrix includes:
[0047] In the case where the data matrix in the next execution stage is the left matrix of the multiplication operation, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices respectively corresponding to each artificial intelligence chip with the column division dimension; or, in the case where the data matrix in the next execution stage is the right matrix of the multiplication operation, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices respectively corresponding to each artificial intelligence chip with the row division dimension.
[0048] In one embodiment, the parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0049] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0050] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, divide the operation result of the artificial intelligence chip in the execution stage into multiple third data sub-matrices corresponding to each artificial intelligence chip respectively with columns as the division dimension according to the number of artificial intelligence chips, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation result of the artificial intelligence chip in the execution stage in the memory resources of the artificial intelligence chip itself.
[0051] In one embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0052] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any one of the artificial intelligence chips, write the operation result of the artificial intelligence chip in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0053] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, divide the operation result of the artificial intelligence chip in the execution stage into multiple fourth data sub-matrices corresponding to each artificial intelligence chip respectively with rows as the division dimension according to the number of artificial intelligence chips, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0054] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation result of the artificial intelligence chip in the execution stage in the memory resources of the artificial intelligence chip itself; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, divide the operation result of the artificial intelligence chip in the execution stage into multiple fifth data sub-matrices corresponding to each artificial intelligence chip respectively with rows as the division dimension according to the number of artificial intelligence chips, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0055] In one embodiment, the parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resource corresponding to the target artificial intelligence chip includes:
[0056] Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction process results:
[0057] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, for any artificial intelligence chip, retain the full reduction process result in the memory resource of the artificial intelligence chip itself;
[0058] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any artificial intelligence chip, divide the full reduction process result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip with the row division dimension according to the number of artificial intelligence chips, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip.
[0059] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the parallel operation method of the operator stream in any one of the above.
[0060] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the parallel operation method of the operator stream in any one of the above.
[0061] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the parallel operation method of the operator stream in any one of the above.
[0062] The above parallel operation method, computer device, and readable storage medium of the operator flow can divide the target operator flow into multiple execution stages. Each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy. The parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy. Further, during the operation of the target operator flow, for any execution stage, the corresponding parallel strategy of the execution stage is adopted inside each artificial intelligence chip to execute the operations corresponding to the operators in the execution stage. And after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, part or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip. The target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips. After the calculations are completed in each execution stage, the operation result of the target operator flow is obtained based on the final operation results of each artificial intelligence chip. By using the parallel operation method, computer device, and readable storage medium of the operator flow provided in the embodiments of the present application, during the execution process of the target operator flow, the corresponding parallel strategy can be accurately adapted according to the operator characteristics of the operators in the target operator flow. At the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips is finely optimized, so that the operators can be fully utilized, the calculation speed of the target operator flow can be significantly improved, and further the inference latency of the large model can be reduced and the inference efficiency of the large model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0064] Figure 1 It is a schematic diagram of the attention operator flow in an embodiment;
[0065] Figure 2 It is a schematic flowchart of the parallel operation method of the operator flow in an embodiment;
[0066] Figure 3 It is a schematic diagram of the parallel processing of the attention operator flow in an embodiment;
[0067] Figure 4 It is a schematic diagram of a computing cluster in another embodiment;
[0068] Figure 5 It is a structural block diagram of the parallel operation device of the operator flow in an embodiment;
[0069] Figure 6 It is the internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0070] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0071] The large model inference process includes multiple computing links, and there may be obvious differences in the matrix computing scale and characteristics involved in each link. Taking the attention operator stream (i.e., the attention operator stream) in the large model inference process as an example, the embodiments of the present application will be described below. It should be understood that the attention operator stream is only an example of the target operator stream in the embodiments of the present application, and is not understood as a limitation on the target operator stream. In fact, any operator stream with matrix scale and characteristic transformation in the computing link is applicable to the embodiments of the present application.
[0072] The attention operator stream is composed of a first MMA (Matrix Multiply - Accumulate) operator, a split - root mean square normalization operator (which can also be denoted as the split_rmsnorm operator, an operator that combines the operations of splitting and root mean square normalization), a second MMA operator, a rotary position embedding operator (which can also be denoted as ROPE (Rotary Position Embedding)), a third MMA operator, a normalized exponential function operator (i.e., the softmax operator), a fourth MMA operator, and an attention multiplication matrix (which can also be denoted as attn_MMA) operator.
[0073] For the attention operator stream, if a single parallel strategy is adopted, it will lead to the problem that some operators are not fully utilized, resulting in waste of operator resources. Taking the tensor parallel strategy as an example, referring to Figure 1 as shown, a schematic diagram of an attention operator stream performing operations using the tensor parallel strategy is shown.
[0074] It should be noted that the attention operator and the shape expressions of different tensors in the embodiments of the present application are only used as an example implementation in the embodiments of the present application, and are not understood as a limitation on the attention operator and the shape expressions of tensors. In fact, for attention operators with different structures, there is a corresponding relationship between the shape expressions of their tensors and the parameters of the attention operator, which is not specifically limited in the embodiments of the present application.
[0075] Referring to Figure 1 the example shown below. First, in the input stage, the input data size is (B, 7168), where B represents the batch size. Enter the first MMA operator. In the current operation stage, the weight matrix size is [7168, 1536 + 576], and matrix multiplication is performed on the input data to obtain an output tensor qc with a shape of (B, 1536 + 576).
[0076] In the normalization and splitting stage, the output tensor qc enters the split-root mean square normalization operator, and the output tensor qc is split into two parts. One part is a tensor q_c with a shape of (B, 1536), and the other part is a tensor kv_c with a shape of (B, 576).
[0077] In the feature processing stage, in this example, it is assumed that the cluster includes 4 GPUs (Graphics Processing Units). The tensor q_c enters the second MMA operator. Since the current tensor parallel strategy is adopted, in the current operation stage, the weight matrix size is [1536, 128×576], so the part allocated to each GPU is a weight matrix with a size of [1536, 32×576]. Matrix multiplication is performed based on this weight matrix, and then through the rotary position embedding operator, only partial data is subjected to rotary encoding processing to obtain a tensor q with a shape of (B×32, 1, 576). It should be noted that Figure 1 for the 4 GPUs in this operation stage, only one of them is taken as an example to show the operation part it executes, and the other 3 GPUs are not shown. Refer to the operation part executed by the shown GPU.
[0078] The tensor kv_c enters the rotary position embedding module, and only partial data among 576 dimensions is subjected to rotary encoding processing to obtain a tensor kv with a shape of (B, S, 576), where S is the sequence length.
[0079] In the attention calculation stage, tensor q and tensor kv enter the third MMA operator together for matrix multiplication, resulting in a tensor s1 with a shape of (B×32, 1, S). Tensor s1 enters the normalized exponential function operator for normalization, and the output tensor s2 still has a shape of (B×32, 1, S), which is used to calculate the attention weights. After the normalization operation, s2 and tensor kv enter the fourth MMA operator together for matrix multiplication, obtaining a tensor O with a shape of (B, 32×512).
[0080] Finally, tensor O enters the attention multiplication matrix operator. The weight matrix shape in the current operation stage is [32×512, 7168]. After matrix multiplication, the final output data with a shape of (B, 7168) is obtained, completing the entire calculation process.
[0081] It can be seen that the tensor parallel strategy can make full use of the second MMA operator and the attention multiplication matrix operator. However, for other computing powers, such as the third MMA operator and the fourth MMA operator, they are not fully utilized, resulting in idle computing power. This situation may affect the low-latency performance of large model inference, making it difficult to meet strict latency requirements.
[0082] The embodiment of the present application provides a method for parallel running of an operator flow. In the execution process of the target operator flow, by dividing the target operator flow into multiple execution stages according to the operator characteristics of each operator in the target operator flow, and precisely matching a corresponding parallel strategy for each execution stage. At the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips can be finely optimized, so that each operator in the target operator flow can be fully utilized, significantly improving the calculation speed of the target operator flow, thereby reducing the large model inference latency and improving the inference efficiency of the large model.
[0083] In an exemplary embodiment, as Figure 2 shown, a method for parallel running of an operator flow is provided. Taking the host side as an example for illustration, it can be understood that the host side may include a CPU (Central Processing Unit). The method includes steps 202 to 206. Among them:
[0084] Step 202: Divide the target operator flow into multiple execution stages. Each execution stage includes at least one operator. The operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy.
[0085] In the embodiments of the present application, the target operator stream can be divided into multiple execution stages based on the matrix calculation scale and characteristics of different operators in the target operator stream, as well as the dependency relationships between the operators. Among them, the principle of division is to make the operators within each execution stage have similarities in the requirements for computing resources and computing characteristics, so as to adopt the same parallel strategy for optimization.
[0086] Exemplarily, starting from the starting operator of the target operator stream, it can be used as the starting operator of the first execution stage. Further, according to the dependency relationships between the operators, the subsequent operators are sequentially added to the current execution stage until an operator is encountered whose matrix calculation scale or calculation characteristics are significantly different from those of the operators within the current stage, or there are complex dependency relationships between this operator and the operators within the current stage, and the same parallel strategy cannot be adopted. At this time, this operator is used as the starting operator of the next execution stage, and the above process is repeated until all the operators in the target operator stream are divided into the corresponding execution stages.
[0087] In an exemplary embodiment, the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy, where: the conventional parallel strategy includes each artificial intelligence chip performing the operation corresponding to the operator on the full data matrix and the full weight matrix.
[0088] The hybrid parallel strategy includes, when the data matrix is the left matrix in a multiplication operation, splitting the data matrix into multiple data sub-matrices with columns as the splitting dimension and splitting the weight matrix into multiple weight sub-matrices with rows as the splitting dimension; or, when the data matrix is the right matrix in a multiplication operation, splitting the data matrix into multiple data sub-matrices with rows as the splitting dimension and splitting the weight matrix into multiple weight sub-matrices with columns as the splitting dimension; each artificial intelligence chip performs the operation corresponding to the operator based on the allocated data sub-matrix and weight sub-matrix.
[0089] In the embodiments of the present application, the artificial intelligence chip includes chips such as GPU (Graphics Processing Unit), GPGPU (General-Purpose Computing on Graphics Processing Units), TPU (Tensor Processing Unit), and NPU (Neural network Processing Unit). In the following embodiments of the present application, the artificial intelligence chip will be taken as an example of GPU for illustration. Taking a cluster composed of 4 GPUs as an example, the conventional parallel strategy, the data parallel strategy, the tensor parallel strategy, and the hybrid parallel strategy are described as follows:
[0090] Conventional parallel strategy: Without any partitioning of the data matrix, the complete data matrix is directly transmitted to 4 GPUs respectively. Each GPU uses the same unpartitioned weight matrix and performs the same arithmetic operation on the received data matrix. Finally, the 4 GPUs output the same arithmetic result.
[0091] Data Parallelism (DP): Partition the data matrix. Split the complete data matrix into 4 parts with approximately the same amount of data in each part, and then write these 4 parts of the data matrix into 4 GPUs respectively. The weight matrix is not partitioned, and the same weight matrix is stored in 4 GPUs respectively. Each GPU performs arithmetic operations on the received partial data matrix and the same weight matrix to obtain its respective partial arithmetic result. Finally, after these partial arithmetic results are aggregated and combined, the final complete arithmetic result is obtained.
[0092] Tensor Parallelism (TP): Partition the weight matrix. For example, partition it into 4 parts by columns with approximately the same amount of data in each part, and then write these 4 parts of the weight matrix into 4 GPUs respectively. The data matrix is not partitioned, and the complete data matrix is written into 4 GPUs respectively. Each GPU performs arithmetic operations on the partial weight matrix it stores and the complete data matrix. Finally, after the arithmetic results of each GPU are aggregated and combined, the complete arithmetic result can be obtained.
[0093] Hybrid parallel strategy: When the data matrix is the left matrix in a multiplication operation, that is, when the data matrix is the left matrix and the weight matrix is the right matrix, split the data matrix by columns as the splitting dimension into 4 data sub-matrices with approximately the same amount of data; at the same time, split the weight matrix by rows as the splitting dimension into 4 weight sub-matrices. Subsequently, distribute these 4 data sub-matrices and 4 weight sub-matrices to 4 GPUs respectively. Or, when the data matrix is the right matrix in a multiplication operation, that is, when the data matrix is the right matrix and the weight matrix is the left matrix, split the data matrix by rows as the splitting dimension into 4 data sub-matrices, and split the weight by columns as the splitting dimension into 4 weight sub-matrices. Each GPU performs arithmetic operations on the received data sub-matrix and weight sub-matrix to complete the calculation of its respective responsible part. Finally, aggregate and integrate the arithmetic results of the 4 GPUs to obtain the final complete arithmetic result.
[0094] It should be noted that the splitting process of the above data matrix and weight matrix can be evenly split, that is, the sizes of the split sub-matrices are exactly the same, or it can be split based on the performance resources of the GPU, that is, the sizes of the split sub-matrices are not exactly the same and are adapted to the performance resources of their respective GPUs. That is, a relatively larger sub-matrix can be divided for a GPU with better performance resources, and a relatively smaller sub-matrix can be divided for a GPU with poorer performance resources, so as to make full use of the GPU resources.
[0095] In the embodiments of the present application, after dividing the target operator stream into multiple execution stages, a corresponding parallel strategy can be matched for each execution stage. In an exemplary embodiment, the execution stage may include at least one of a first execution stage, a second execution stage, a third execution stage, and a fourth execution stage, where:
[0096] The first execution stage includes at least one first operator, and the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirement. The first execution stage corresponds to a conventional parallel strategy;
[0097] The second execution stage includes at least one second operator, and the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirement, but the weight matrix of the operation satisfies the parallel requirement. The second execution stage corresponds to a tensor parallel strategy;
[0098] The third execution stage includes at least one third operator, and the third operator is an operator for which the data matrix of the operation satisfies the parallel requirement, but the weight matrix of the operation does not satisfy the parallel requirement. The third execution stage corresponds to a data parallel strategy;
[0099] The fourth execution stage includes at least one fourth operator, and the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirement. The fourth execution stage corresponds to a hybrid parallel strategy.
[0100] In the embodiments of the present application, in the first execution stage, there is at least one first operator. The data matrix and weight matrix of such a first operator are both relatively small. Due to the limited matrix size, it is difficult to significantly improve the efficiency through parallel computing, that is, neither of them satisfies the parallel computing conditions. Based on this characteristic, it can be determined that the first execution stage adopts a conventional parallel strategy. Under this conventional parallel strategy, the complete data matrix and weight matrix will be written to each GPU respectively, and each GPU will perform the same full-scale operation, and finally output the same calculation result.
[0101] In the second execution phase, there is at least one second operator. Such second operators exhibit the characteristics of a relatively small data matrix size and a relatively large weight matrix size. Specifically, due to the small size of its data matrix, it cannot fully utilize parallel computing resources to improve efficiency and does not meet the requirements of parallel computing; while the weight matrix, because of its large size, has the conditions for parallel computing. Based on this characteristic, it can be determined that the second execution phase adopts a tensor parallel strategy. Under this tensor parallel strategy, each GPU divides the weight matrix and stores it distributively. At the same time, each GPU receives the complete data matrix, and each GPU uses the stored partial weight matrix and the complete data matrix for operations, and can efficiently complete the operation processing tasks of the second operator.
[0102] In the third execution phase, there is at least one third operator. Such third operators exhibit the characteristics of a relatively large data matrix size and a relatively small weight matrix size. Specifically, due to the small size of its weight matrix, it cannot fully utilize parallel computing resources to improve efficiency and does not meet the requirements of parallel computing; while the data matrix, because of its large size, has the conditions for parallel computing. Based on this characteristic, it can be determined that the third execution phase adopts a data parallel strategy. Under this data parallel strategy, the data matrix is split into multiple parts and distributed to each GPU respectively; while the complete weight matrix is copied to each GPU. Each GPU uses the allocated partial data matrix and the complete weight matrix for operations, thereby efficiently completing the operation processing of the third operator.
[0103] In the fourth execution phase, there is at least one fourth operator. Such fourth operators exhibit the characteristics of both a relatively large data matrix size and a relatively large weight matrix size. Specifically, due to their large sizes, both the weight matrix and the data matrix have the conditions for parallel computing. Based on this characteristic, it can be determined that the fourth execution phase adopts a hybrid parallel strategy. Under this hybrid parallel strategy, each GPU divides the weight matrix and stores it distributively, and the data matrix is split into multiple parts and distributed to each GPU respectively. Each GPU uses the allocated partial data matrix and partial weight matrix for operations, thereby efficiently completing the operation processing of the fourth operator.
[0104] Step 204, during the running of the target operator stream, for any execution phase, within each artificial intelligence chip, adopt the parallel strategy corresponding to the execution phase to execute the operations corresponding to each operator within the execution phase. And after the operations in the execution phase are completed, based on the parallel strategy corresponding to the next execution phase, directly write some or all of the operation results of each artificial intelligence chip in the execution phase into the memory resources corresponding to the target artificial intelligence chip, where the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips.
[0105] In the embodiments of the present application, when the target operator flow is in the running (i.e., computing) state, starting from the starting operator, operations will be carried out according to the parallel strategy corresponding to its execution stage. After the operations in the current execution stage are completed, based on the parallel strategy of the next execution stage, part or all of the content of the GPU operation results in the current execution stage will be directly written to the memory resources of the target GPU. Here, the direct writing operation means that only the data required for the operations of the target GPU in the next execution stage is transmitted through the communication link between GPUs. The embodiments of the present application do not specifically limit the communication method between GPUs. For example, if there is a P2P link (Peer-to-Peer Link) connection between GPUs, communication can be carried out through this P2P link connection; if there is no P2P link connection, data communication can also be achieved using the PCIE (Peripheral Component Interconnect Express, high-speed peripheral component interconnect standard) bus.
[0106] In this way, the amount of data transferred between GPUs can be effectively reduced, the communication time consumption can be reduced, and thus the inference latency can be reduced, significantly improving the inference efficiency of the large model. The specific process of the direct writing operation will be elaborated in detail in the following embodiments, and the embodiments of the present application will not describe it in detail here.
[0107] Step 206, after the calculations are completed in each execution stage, the operation result of the target operator flow is obtained based on the final operation results of each artificial intelligence chip.
[0108] In the embodiments of the present application, after the calculations are completed in each execution stage, that is, after each GPU has completed the operation of the target operator flow and obtained the final operation result, the operation result of the target operator flow can be obtained by summarizing and merging the final operation results of each GPU. The specific summarizing and merging method can be determined based on the parallel strategy adopted in the last stage of each GPU. The embodiments of the present application will not elaborate on the summarizing and merging method.
[0109] By adopting the parallel operation method of the operator flow provided by the embodiments of the present application, during the execution process of the target operator flow, the corresponding parallel strategy can be accurately adapted according to the operator characteristics of each operator in the target operator flow. At the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips is finely optimized, so that the operators can be fully utilized, significantly improving the calculation speed of the target operator flow, thereby reducing the inference latency of the large model and improving the inference efficiency of the large model.
[0110] In an exemplary embodiment, the parallel strategy corresponding to the execution stage is a conventional parallel strategy. Based on the parallel strategy corresponding to the next execution stage, some or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip, including at least one of the following:
[0111] If the parallel strategy corresponding to the next execution stage is a tensor parallel strategy, for any artificial intelligence chip, the operation result of the artificial intelligence chip in the execution stage is retained in the memory resources of the artificial intelligence chip itself;
[0112] If the parallel strategy corresponding to the next execution stage is a data parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, with the row division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into multiple first data sub-matrices corresponding to each artificial intelligence chip respectively, and the first data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resources of the artificial intelligence chip itself;
[0113] If the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is divided into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively, and the second data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resources of the artificial intelligence chip itself, where the division method of the second data sub-matrix includes:
[0114] In the case where the data matrix in the next execution stage is the left matrix of the multiplication operation, with the column division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively; or, in the case where the data matrix in the next execution stage is the right matrix of the multiplication operation, with the row division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively.
[0115] In the embodiment of the present application, it is assumed that the current execution stage corresponds to the first execution stage, and the corresponding parallel strategy is a conventional parallel strategy. The conventional parallel strategy refers to the operation between the full data matrix A (scale (m, k)) and the full weight matrix B (scale (k, n)). After the operation is completed, each GPU will obtain the same operation result, and this operation result is a complete data matrix C (scale (m, n)), which is the input data for the next execution stage.
[0116] In one example, the parallel strategy corresponding to the next execution stage is the tensor parallel strategy. Since the tensor parallel strategy uses the complete data matrix to perform operations with the partial weight matrix stored inside the GPU, at this time, each GPU can calculate and obtain the complete data matrix C without data synchronization with other GPUs. Therefore, each GPU can directly write the operation result into its own memory resource at this time.
[0117] In another example, the parallel strategy corresponding to the next execution stage is the data parallel strategy. The data parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs, and then writes the multiple data matrices into each GPU to perform operations with the full weight matrix. Since each GPU calculates and obtains the complete data matrix at this time, the operation result can be directly divided into multiple first data sub-matrices with the row as the division dimension. Each first data sub-matrix corresponds to one GPU, and the first data sub-matrix corresponding to each GPU itself is retained in the memory resource of the GPU itself.
[0118] Exemplarily, assume that there are 4 GPUs in the cluster, namely GPU0, GPU1, GPU2, and GPU3. The operation result obtained by each GPU in the current execution stage is a data matrix C1 with a size of (m, n). If the data parallel strategy is adopted in the next execution stage, each GPU should operate on a partial data matrix with a size of (m / 4, n). Then, the division operation with the row as the division dimension can be performed on the operation result data matrix C1 of each GPU to obtain 4 first data sub-matrices with a size of (m / 4, n):
[0119] The first data sub-matrix C 10 : corresponding to the 1st to m / 4th rows of C1;
[0120] The first data sub-matrix C 11 : corresponding to the (m / 4 + 1)th to m / 2th rows of C1;
[0121] The first data sub-matrix C 12 : corresponding to the (m / 2 + 1)th to 3m / 4th rows of C1;
[0122] The first data sub-matrix C 13 : corresponding to the (3m / 4 + 1)th to mth rows of C1.
[0123] Furthermore, C 10 can be retained in the memory resource of GPU0 itself, C 11 can be retained in the memory resource of GPU1 itself, C 12 can be retained in the memory resource of GPU2 itself, C 13It is stored in the memory resources of the GPU 3 itself. In the operation process of the next execution stage, the first data sub-matrix written and the full weight matrix B2 (with a size of (n, p)) are respectively operated inside each GPU to obtain an operation result C with a size of (m / 4, p). 20 and C 21 and C 22 and C 23 . Among them, C 20 corresponds to the 1st to m / 4th rows of the final operation result C2, C 21 corresponds to the (m / 4 + 1)th to m / 2th rows of the final operation result C2, C 22 corresponds to the (m / 2 + 1)th to 3m / 4th rows of the final operation result C2, C 23 corresponds to the (3m / 4 + 1)th to mth rows of the final operation result C2. C 20 and C 21 and C 22 and C 23 After summarizing and merging, it is the final operation result C2 (with a size of (m, p)) of the operator in the next execution stage.
[0124] In another example, the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy. The hybrid parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs with columns or rows as the division dimension, and divides the complete weight matrix into multiple weight matrices with rows or columns as the division dimension and stores them in each GPU. Then, the multiple data matrices are respectively written into each GPU to operate with the partial weight matrix stored in it. Since each GPU has calculated the complete data matrix at this time, the operation result can be directly divided into multiple second data sub-matrices with columns or rows as the division dimension. Each second data sub-matrix corresponds to one GPU, and the second data sub-matrix corresponding to each GPU itself is stored in the memory resources of the GPU itself.
[0125] In one example, in the multiplication operation process of the next execution stage, the data matrix is the left matrix and the weight matrix is the right matrix. Exemplarily, still taking the 4 GPUs in the previous example as an example, the operation result obtained by each GPU in the current operation stage is a data matrix C3 with a size of (m, k). Then, each GPU can perform the division operation on the data matrix C3 with columns as the splitting dimension to obtain 4 second data sub-matrices with a size of (m, k / 4):
[0126] Second data sub-matrix C 30 : corresponding to the 1st to k / 4th columns of C3;
[0127] Second data sub-matrix C 31 : corresponding to the (k / 4 + 1)th to k / 2th columns of C3;
[0128] The second data sub-matrix C 32 : corresponding to the columns k / 2 + 1 to 3k / 4 of C3;
[0129] The second data sub-matrix C 33 : corresponding to the columns 3k / 4 + 1 to k of C3.
[0130] Meanwhile, for the weight matrix B3 (with size (k, n)) in the current operation stage, a partitioning operation is also performed with the row as the splitting dimension, resulting in 4 weight sub-matrices with size (k / 4, n):
[0131] The weight sub-matrix B stored in GPU0 30 : corresponding to the rows 1 to k / 4 of B3;
[0132] The weight sub-matrix B stored in GPU1 31 : corresponding to the rows k / 4 + 1 to k / 2 of B3;
[0133] The weight sub-matrix B stored in GPU2 32 : corresponding to the rows k / 2 + 1 to 3k / 4 of B3;
[0134] The weight sub-matrix B stored in GPU3 33 : corresponding to the rows 3k / 4 + 1 to k of B3.
[0135] Alternatively, in another example, during the multiplication operation in the next execution stage, the data matrix is the right matrix and the weight matrix is the left matrix. Exemplarily, still taking the 4 GPUs in the previous example, the operation result obtained by each GPU in the current operation stage is a data matrix C3 with size (k, n). Then each GPU can perform a partitioning operation on the data matrix C3 with the row as the splitting dimension, resulting in 4 second data sub-matrices with size (k / 4, n):
[0136] The second data sub-matrix C 30 : corresponding to the rows 1 to k / 4 of C3;
[0137] The second data sub-matrix C 31 : corresponding to the rows k / 4 + 1 to k / 2 of C3;
[0138] The second data sub-matrix C 32 : corresponding to the rows k / 2 + 1 to 3k / 4 of C3;
[0139] The second data sub-matrix C 33 : corresponding to the rows 3k / 4 + 1 to k of C3.
[0140] Meanwhile, for the weight matrix B3 (with size (m, k)) in the current operation stage, a partitioning operation is also performed with columns as the partitioning dimension, resulting in 4 weight sub-matrices with size (m, k / 4):
[0141] The weight sub-matrix B stored in GPU0 30 : corresponding to the 1st to k / 4th columns of B3;
[0142] The weight sub-matrix B stored in GPU1 31 : corresponding to the (k / 4 + 1)th to k / 2th columns of B3;
[0143] The weight sub-matrix B stored in GPU2 32 : corresponding to the (k / 2 + 1)th to 3k / 4th columns of B3;
[0144] The weight sub-matrix B stored in GPU3 33 : corresponding to the (3k / 4 + 1)th to kth columns of B3.
[0145] Furthermore, C 30 can be retained in the memory resources of GPU0 itself. In the next execution stage, the operation with B 30 is performed within GPU0, and C 31 is retained in the memory resources of GPU1 itself. In the next execution stage, the operation with B 31 is performed within GPU1, C 32 is retained in the memory resources of GPU2 itself. In the next execution stage, the operation with B 32 is performed within GPU2, C 33 is retained in the memory resources of GPU3 itself. In the next execution stage, the operation with B 33 is performed within GPU3, and then the operation results C 40 、C 41 、C 42 and C 43 with size (m, n) are obtained respectively. Among them, C 40 、C 41 、C 42 and C 43 are aggregated and combined, which is the final operation result C4 of the operator in the next execution stage.
[0146] In this way, in the case of the conventional parallel strategy corresponding to the current execution stage, based on the parallel strategy corresponding to the next execution stage, the data required for the operator operation in the next execution stage can be directly written into the memory resources of each GPU, thereby reducing the amount of data transferred between GPUs, reducing the communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0147] In an exemplary embodiment, the parallel strategy corresponding to the execution stage is a data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, part or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip, including at least one of the following:
[0148] If the parallel strategy corresponding to the next execution stage is a conventional parallel strategy or a tensor parallel strategy, for any artificial intelligence chip, the operation results of the artificial intelligence chip in the execution stage are written into the memory resources of each artificial intelligence chip respectively;
[0149] If the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, with the column as the division dimension, the operation results of the artificial intelligence chip in the execution stage are divided into multiple third data sub-matrices corresponding to each artificial intelligence chip respectively, and the multiple third data sub-matrices are written into the memory resources corresponding to their respective artificial intelligence chips respectively; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, the operation results of the artificial intelligence chip in the execution stage are retained in the memory resources of the artificial intelligence chip itself.
[0150] In the embodiment of the present application, it is assumed that the current execution stage corresponds to the third execution stage, and its corresponding parallel strategy is a data parallel strategy. The data parallel strategy means that the data matrix A is divided into multiple parts and written into the memory resources of the corresponding GPUs respectively, and in each GPU, it performs operations with the full-weight matrix B1. After the operation is completed, the operation results of each GPU are part of the operation results of the operator in the current execution stage. By summarizing and combining the operation results of each GPU, the operation results of the operator in the current execution stage can be obtained.
[0151] Exemplarily, in the previous example, GPU0, GPU1, GPU2, and GPU3 respectively obtain the operation results C 20 、C 21 、C 22 and C 23 as an example.
[0152] In an example, the parallel strategy corresponding to the next execution stage is a tensor parallel strategy or a conventional parallel strategy. Since both the tensor parallel strategy and the conventional parallel strategy use the complete data matrix to perform operations with part or all of the weight matrix stored inside the GPU, and the operation results of each GPU in the current execution stage (data parallel stage) are part of the complete operation results, data synchronization is required between each GPU at this time to write their respective operation results into their own memory resources and the memory resources of other GPUs.
[0153] For example, C 20 is retained in the memory resources of GPU0 and written into the memory resources of GPU1, GPU2, and GPU3 respectively. C 21 is retained in the memory resources of GPU1 and written into the memory resources of GPU0, GPU2, and GPU3 respectively. C 22 is retained in the memory resources of GPU2 and written into the memory resources of GPU0, GPU1, and GPU3 respectively. C 23 is retained in the memory resources of GPU3 and written into the memory resources of GPU0, GPU1, and GPU2 respectively. In this way, each GPU can obtain C 20 , C 21 , C 22 , and C 23 . After summarizing and merging, the complete operation result C2 can be obtained. Then, after entering the next execution stage, it can perform operations with some or all of the weight matrices stored inside each GPU to complete tensor parallel or regular parallel operations and obtain the corresponding operation results.
[0154] In another example, the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy. The hybrid parallel strategy divides the complete data matrix into multiple data matrices by columns or rows according to the number of GPUs, and divides the complete weight matrix into multiple weight matrices by rows or columns and stores them in each GPU. For example, the weight matrix W with a shape of (k, p) is evenly divided into 4 weight sub-matrices W0, W1, W2, and W3 by rows, and each sub-matrix has a shape of (k / 4, p), which are respectively stored in each GPU. Then, the multiple data matrices divided by columns are respectively written into the memory resources of each GPU to perform operations with the partial weight matrices stored therein; or, the weight matrix W with a shape of (p, k) is evenly divided into 4 weight sub-matrices W0, W1, W2, and W3 by columns, and each sub-matrix has a shape of (p, k / 4), which are respectively stored in each GPU. Then, the multiple data matrices divided by rows are respectively written into the memory resources of each GPU to perform operations with the partial weight matrices stored therein.
[0155] Since after the operations in the current execution stage (data parallel stage) are completed, the operation results of each GPU respectively correspond to partial rows of the complete operation result, and the final operation result of the operator in the current execution stage is obtained only after merging. When entering the next execution stage, since the next execution stage is a hybrid parallel strategy, if the data matrix in the next execution stage is the left matrix and the weight matrix is the right matrix, then each GPU will divide and obtain partial columns of the final operation result. Taking the scale of the operation result of each GPU as (m / 4, k) and the final operation result as (m, k) as an example, each GPU in the next execution stage will divide and obtain a data matrix with a scale of (m, k / 4).
[0156] Therefore, after the operation in the current execution stage is completed, after dividing the operation results of each GPU by column to obtain multiple third data sub-matrices, each third data sub-matrix can be written into the memory resources of each GPU respectively, and spliced by row. Each GPU can obtain partial columns of the final operation result in the current execution stage, so as to enter the next execution stage. In each GPU, it can directly perform operations with the partial weight matrix obtained by row division stored, and obtain the corresponding operation result.
[0157] Taking the 4 GPUs in the previous example as an example, the operation result C of each GPU 20 、C 21 、C 22 and C 23 all perform the operation of evenly dividing by column into 4 third data sub-matrices, and obtain the third data sub-matrices G i0 、G i1 、G i2 、G i3 , where i is the index of the GPU (the value range is 0, 1, 2, 3), and the scale of each third data sub-matrix is (m / 4, k / 4), where:
[0158] The third data sub-matrix G i0 corresponds to the 1st to k / 4th columns, the third data sub-matrix G i1 corresponds to the (k / 4 + 1)th to k / 2th columns, the third data sub-matrix G i2 corresponds to the (k / 2 + 1)th to 3k / 4th columns, and the third data sub-matrix G i3 corresponds to the (3k / 4 + 1)th to kth columns.
[0159] Directly write G ij into the memory resource of GPU j , then GPU j will obtain 4 third data sub-matrices with a scale of (m / 4, k / 4). After splicing by row, a data sub-matrix with a scale of (m, k / 4) can be obtained, where j is the index of the GPU (the value range is 0, 1, 2, 3). When entering the operation of the next execution stage, the spliced data sub-matrix is operated with the weight sub-matrix obtained by row division inside the GPU, and an operation result with a scale of (m, p) can be obtained. By performing allreduce processing on the operation results of each GPU, the final operation result corresponding to the operator in this operation stage can be obtained.
[0160] Alternatively, if the data matrix in the next execution stage is the right matrix and the weight matrix is the left matrix, each GPU will divide and obtain partial rows of the final operation result. Taking the operation result scale of each GPU as (k / 4, m) and the final operation result as (k, m) as an example, each GPU in the next execution stage will divide and obtain a data matrix with a scale of (k / 4, m). Therefore, after the operation in the current execution stage is completed, the operation results of each GPU can be retained in the memory resources of each GPU itself.
[0161] Exemplarily, taking GPU0, GPU1, GPU2, and GPU3 as examples to obtain operation results C with a scale of (k / 4, m) respectively 20 、C 21 、C 22 and C 23 as an example, then C 20 can be retained in the memory resource of GPU0 and operated with the weight sub-matrix W0 with a scale of (p, k / 4) stored in GPU0; C 21 can be retained in the memory resource of GPU1 and operated with the weight sub-matrix W1 with a scale of (p, k / 4) stored in GPU1; C 22 can be retained in the memory resource of GPU2 and operated with the weight sub-matrix W2 with a scale of (p, k / 4) stored in GPU2; C 23 can be retained in the memory resource of GPU3 and operated with the weight sub-matrix W3 with a scale of (p, k / 4) stored in GPU3. In this way, each GPU can obtain partial rows of the final operation result in the current execution stage, and thus enter the next execution stage, where it can directly operate with the partial weight matrix obtained by column division stored in each GPU to obtain the corresponding operation result.
[0162] The method provided in the embodiments of the present application, in the case of corresponding data parallel strategies in the current execution stage, can directly write the data required for the operator operation in the next execution stage into the memory resources of each GPU based on the corresponding parallel strategy in the next execution stage, thereby reducing the amount of data transmitted between GPUs, reducing communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0163] In an exemplary embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the corresponding parallel strategy in the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0164] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any artificial intelligence chip, write the operation results of the artificial intelligence chip in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0165] If the parallel strategy corresponding to the next execution stage is a data parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, with the row division dimension, divide the operation results of the artificial intelligence chip in the execution stage into multiple fourth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0166] If the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy, for any artificial intelligence chip, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chip in the execution stage in the memory resources of the artificial intelligence chip itself; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with the row division dimension, divide the operation results of the artificial intelligence chip in the execution stage into multiple fifth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0167] In the embodiments of the present application, it is assumed that the current execution stage corresponds to the second execution stage, and the corresponding parallel strategy is a tensor parallel strategy. The tensor parallel strategy means that the weight matrix B is divided into multiple parts and stored in the corresponding GPUs respectively. Inside each GPU, it performs operations with the full data matrix. After the operation is completed, the operation results of each GPU are part of the operation results of the operator in the current execution stage. By aggregating and combining the operation results of each GPU, the operation results of the operator in the current execution stage can be obtained.
[0168] Exemplarily, still taking the 4 GPUs, GPU0, GPU1, GPU2, and GPU3 in the previous example as an example, in the current execution stage (tensor parallel stage), the operation results obtained respectively are C with a scale of (m, q / 4) 50 、C 51 、C 52 and C 53 . Taking C 50 、C 51 、C 52 and C 53 as an example, aggregating and combining C
[0169] In one example, the parallel strategy corresponding to the next execution phase is a conventional parallel strategy. Since the conventional parallel strategy performs operations using the complete data matrix and the complete weight matrix, and the operation results of each GPU in the current execution phase are part of the complete operation result C5, data synchronization is required among the GPUs at this time to write their respective operation results into their own memory resources and the memory resources of other GPUs. In this way, each GPU can obtain the operation result C 50 、C 51 、C 52 and C 53 . After summarizing and merging, the complete operation result C5 can be obtained, and then the operation can be performed within each GPU with the complete weight matrix stored therein to obtain the corresponding operation result.
[0170] In another example, the parallel strategy corresponding to the next execution phase is a data parallel strategy. The data parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs, and then writes the multiple data matrices into the memory resources of each GPU respectively for operation with the full weight matrix. That is, in the next execution phase, each GPU operates on a partial row of the complete operation result of the operator in the current execution phase.
[0171] Since the operation result obtained by each GPU in the current execution phase (tensor parallel phase) is a partial column of the complete operation result of the operator in the current execution phase, the operation results of each GPU can be directly divided into multiple fourth data sub-matrices corresponding to each GPU respectively with the row as the division dimension.
[0172] Taking the 4 GPUs in the aforementioned example as an example, the operation result C 50 、C 51 、C 52 and C 53 of each GPU all perform the operation of evenly dividing into 4 fourth data sub-matrices by row to obtain the fourth data sub-matrix R i0 、R i1 、R i2 、R i3 , where i is the index of the GPU (taking values of 0, 1, 2, 3), and the scale of each fourth data sub-matrix is (m / 4, k / 4), where:
[0173] The fourth data sub-matrix R i0 corresponds to columns 1 to m / 4, the fourth data sub-matrix R i1 corresponds to columns m / 4 + 1 to m / 2, the fourth data sub-matrix R i2 corresponds to columns m / 2 + 1 to 3m / 4, and the fourth data sub-matrix R i3 corresponds to columns 3m / 4 + 1 to m.
[0174] Write R ij directly into the memory resources of the GPU j , then the GPU j will obtain 4 fourth data sub-matrices with dimensions of (m / 4, k / 4). After concatenating them by columns, a data sub-matrix with dimensions of (m / 4, k) can be obtained. Here, j is the index of the GPU (taking values 0, 1, 2, 3). When entering the operations of the next execution stage, the concatenated data sub-matrix is operated on with the complete weight matrix (k, p) stored inside the GPU, and an operation result with dimensions of (m / 4, p) can be obtained. By aggregating and combining the operation results of each GPU, the final operation result with dimensions of (m, p) corresponding to the operator in this running stage can be obtained.
[0175] In another example, the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy. The hybrid parallel strategy is to divide the complete weight matrix into multiple weight matrices and store them in each GPU. For example: when the data matrix is the left matrix and the weight matrix is the right matrix, the weight matrix W with shape (k, p) can be evenly divided into 4 weight sub-matrices W0, W1, W2, W3 by rows, each sub-matrix with shape (k / 4, p), and stored in each GPU respectively. Then, according to the number of GPUs, the complete data matrix is divided into multiple data matrices by columns, and these multiple data matrices are written into the memory resources of each GPU respectively to perform operations with the partial weight matrices stored in them; or, when the data matrix is the right matrix and the weight matrix is the left matrix, the weight matrix W with shape (p, k) can be evenly divided into 4 weight sub-matrices W0, W1, W2, W3 by columns, each sub-matrix with shape (p, k / 4), and stored in each GPU respectively. Then, according to the number of GPUs, the complete data matrix is divided into multiple data matrices by rows, and these multiple data matrices are written into the memory resources of each GPU respectively to perform operations with the partial weight matrices stored in them.
[0176] In one example, the data matrix in the next execution stage is the left matrix for multiplication operations. Since after the current execution stage (tensor parallel stage) is completed, the operation result obtained by each GPU is a partial column in the complete operation result of the operator in the current execution stage, with dimensions of (m, k / 4), then at this time, it is not necessary to further divide this operation result and directly retain it in the memory resources of the GPU itself. In the next execution stage, this operation result is used to perform operations with the weight sub-matrices stored inside the GPU, and an operation result with dimensions of (m, p) can be obtained. By performing an allreduce operation on the current operation results of each GPU with dimensions of (m, p), the final operation result corresponding to the operator in this running stage can be obtained.
[0177] In another example, the data matrix in the next execution stage is the right matrix for the multiplication operation. Since after the current execution stage (tensor parallelism stage) is completed, the operation result obtained by each GPU is a partial column in the complete operation result of the operator in the current execution stage, with a size of (m, k / 4), at this time, the operation result can be further divided by rows, and the fifth data sub-matrices with a size of (m / 4, k / 4) obtained by the division are respectively written into the memory resources of the corresponding GPUs. After each GPU merges the multiple written fifth data sub-matrices by columns, the GPU can obtain a data matrix with a size of (m / 4, k) (i.e., partial rows of the final operation result in the current execution stage). In the next execution stage, this data matrix is used to perform an operation with the weight sub-matrix stored inside the GPU, and the operation results of each GPU are fully reduced to obtain the final operation result corresponding to the operator in this operation stage. The specific process can refer to the relevant description in the foregoing embodiments, and will not be elaborated herein in the embodiments of the present application.
[0178] In this way, in the case of the tensor parallelism strategy corresponding to the current execution stage, based on the parallelism strategy corresponding to the next execution stage, the data required for the operator operation in the next execution stage can be directly written into the memory resources of each GPU, thereby reducing the amount of data transmitted between GPUs, reducing communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0179] In an exemplary embodiment, the parallelism strategy corresponding to the execution stage is a hybrid parallelism strategy. Based on the parallelism strategy corresponding to the next execution stage, part or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip, including:
[0180] Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction process result:
[0181] If the parallelism strategy corresponding to the next execution stage is a tensor parallelism strategy or a conventional parallelism strategy, for any artificial intelligence chip, retain the full reduction process result in the memory resources of the artificial intelligence chip itself;
[0182] If the parallelism strategy corresponding to the next execution stage is a data parallelism strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, with the row as the division dimension, divide the full reduction process result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself.
[0183] In the embodiments of the present application, it is assumed that the current execution stage corresponds to the fourth execution stage, and its corresponding parallel strategy is a hybrid parallel strategy. The hybrid parallel strategy can be referred to the description of the foregoing embodiments, and will not be elaborated herein. After each GPU performs operations using the hybrid parallel strategy, it is necessary to perform a global reduction process on the operation results of each artificial intelligence chip in the execution stage to obtain the global reduction process result.
[0184] In one example, the parallel strategy corresponding to the next execution stage is a tensor parallel strategy or a conventional parallel strategy. Since both the tensor parallel strategy and the conventional parallel strategy operate on a complete data matrix, the global reduction process result is a complete data matrix. Therefore, for any GPU, there is no need to perform data synchronization, and the global reduction process result can be directly retained in the memory resources of the GPU itself and used for operations with the complete or partial weight matrix stored in the GPU.
[0185] In another example, the parallel strategy corresponding to the next execution stage is a data parallel strategy. Since the data parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs, and then writes the multiple data matrices into the memory resources of each GPU respectively for operations with the full amount of the weight matrix. Since the global reduction process result is a complete data matrix, the global reduction process result can be directly divided into multiple sixth data sub-matrices with the row as the division dimension. Each sixth data sub-matrix corresponds to a GPU, and the sixth data sub-matrix corresponding to each GPU itself can be retained in the memory resources of the GPU itself. The specific process can refer to the division and direct writing operation of the foregoing first data sub-matrix.
[0186] In this way, in the case where the current execution stage corresponds to the hybrid parallel strategy, based on the parallel strategy corresponding to the next execution stage, the data required for the operator operation in the next execution stage can be directly written into the memory resources of each GPU, thereby reducing the amount of data transmitted between GPUs, reducing the communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0187] To enable those skilled in the art to better understand the embodiments of the present application, the following will illustrate the embodiments of the present application through specific examples.
[0188] Exemplarily, taking Figure 1 the attention operator stream shown as an example, using the operation method of the operator stream provided by the embodiments of the present application, the operators in the attention operator stream are divided into four execution stages according to the matrix calculation scale and characteristics, as shown in reference to Figure 3 shown.
[0189] Among them, the data matrices and weight matrices of the operations of the first MMA operator, the split-root mean square normalization operator, and the rotary position embedding operator executed by the tensor kv_c are all relatively small, so they are divided into the same execution stage, that is, the first execution stage, and this execution stage corresponds to the conventional parallel strategy. Figure 3 In the above, the operations of the first MMA operator, the split-root mean square normalization operator, and the rotary position embedding operator executed by the tensor kv_c are fused into a matrix multiply-accumulate-root mean square normalization-update cache operator (which can also be expressed as MMA_rmsnorm_updateCache).
[0190] The data matrix of the operation of the second MMA operator and the position embedding operator executed by the tensor q_c is relatively small, but the weight matrix is relatively large, so they are divided into the same execution stage, that is, the second execution stage, and this execution stage corresponds to the tensor parallel strategy. Figure 3 In the above, the second MMA operator and the position embedding operator executed by the tensor q_c are fused into a matrix multiply-accumulate-query embedding operator (which can also be expressed as MMA_query_embedding).
[0191] The data matrices of the operations of the third MMA operator, the exponential normalization function operator, and the fourth MMA operator are relatively large, but the weight matrices are relatively small, so they are divided into the same execution stage, that is, the third execution stage, and this execution stage corresponds to the data parallel strategy. Figure 3 In the above, the third MMA operator, the exponential normalization function operator, and the fourth MMA operator are fused into a grouped query attention operator (Grouped Query Attention, GQA).
[0192] The data matrix and weight matrix of the attention multiplication matrix operator are both relatively large, and it is used as an execution stage, that is, the fourth execution stage, and this execution stage corresponds to the hybrid parallel strategy.
[0193] The input data enters the first execution stage. Since the first execution stage corresponds to the conventional parallel strategy, that is, the input data is written into each GPU, and the operation of the input data and the complete weight matrix is performed in each GPU, obtaining an intermediate tensor with a shape of (B, 1536), and an intermediate tensor with a shape of (B, S, 576) (it should be noted that this intermediate tensor is not shown in Figure 3 ).
[0194] Since the second execution stage adopts the tensor parallel strategy, for the operation results of each GPU in the first execution stage, the operation results (tensor q_c and tensor kv) can be directly retained in the memory resources of each GPU, enter the second execution stage, and perform tensor parallel operations on the written data matrix and a part of the weight matrix inside each GPU.
[0195] After the tensor parallel operation is completed in the second execution stage, each GPU obtains an intermediate tensor with a shape of (B, 128×576 / 4). After the operation results of each GPU are aggregated and combined, the operation result of the operator in the second execution stage is obtained. Since the data parallel strategy is adopted in the third execution stage, for the operation results of each GPU in the second execution stage, the operation results can be split into 4 data matrices corresponding to 4 GPUs respectively according to rows (with a shape of (B / 4, 128×576 / 4)), and the split data matrices are respectively written into the memory resources of the corresponding GPUs ( Figure 3 Only the split schematic of the operation result of one GPU is shown in [[ ]], and the splits of the operation results of other GPUs can be referred to), then at this time each GPU obtains partial rows of the operation result of the operator in the second execution stage and enters the third execution stage to perform data parallel operations on the partial data matrices written in each GPU and the complete weight matrix.
[0196] After the data parallel operation is completed in the third execution stage, each GPU obtains an intermediate tensor with a shape of (B / 4, 128×576). Since the hybrid parallel strategy is adopted in the fourth stage, for the operation results of each GPU in the third execution stage, the operation results can be split into 4 data matrices corresponding to 4 GPUs respectively according to columns (with a shape of (B / 4, 128×512 / 4)), and the split data matrices are respectively written into the memory resources of the corresponding GPUs ( Figure 3 Only the split schematic of the operation result of one GPU is shown in [[ ]], and the splits of the operation results of other GPUs can be referred to), then at this time each GPU obtains partial columns of the operation result of the operator in the third execution stage and enters the fourth execution stage to perform data parallel operations on the partial data matrices written in each GPU and the partial weight matrix stored.
[0197] After the hybrid parallel operation is completed in the fourth execution stage, the operation results output by the four GPUs are allreduced to obtain the final result.
[0198] The parallel operation method of the operator stream provided by the embodiment of the present application can flexibly select a parallel strategy according to the calculation characteristics of the operator stream, and after the calculation is completed, according to the requirements of the subsequent stage, part or all of the operation results are directly written to the memory resources of the corresponding GPU card, giving full play to the operator performance, which can not only meet the better latency performance requirements but also have better throughput.
[0199] In addition, it should be noted that the parallel operation method of the operator flow provided in the embodiments of this application can be applied at least to the inference processes in fields such as speech processing, image processing, text processing, and video processing. Among them, in the field of speech processing, the data matrices of the operations of the operators in the target operator flow can be audio data or audio feature data; in the field of image processing, the data matrices of the operations of the operators in the target operator flow can be image data or image feature data; in the field of text processing, the data matrices of the operations of the operators in the target operator flow can be text data or text feature data; in the field of video processing, the data matrices of the operations of the operators in the target operator flow can be video data or video feature data, etc.
[0200] In the field of image recognition, for complex large model inference tasks, a large amount of computation and data transmission are often involved, especially when dealing with models based on the attention mechanism. The following will elaborate on the application of this solution in the image recognition scenario by combining the use of 4 GPUs to accelerate the inference process of the attention operator flow.
[0201] Suppose a batch of image data needs to be recognized to detect target objects therein, such as: whether there are specific geographical features, such as airports, ports, etc.; or whether there are people or items with specific features. In this example, a large model based on the Transformer architecture is adopted, where the attention mechanism is the core component of the model, and the target operator flow includes the attention mechanism and multiple related operators before and after it.
[0202] First, according to the matrix scale of the operations of the operators in the target operator flow and the dependencies between them, the target operator flow is divided into multiple execution stages, and a suitable parallel strategy is selected for each execution stage.
[0203] For example: The operators for feature extraction and preliminary transformation are divided into the first execution stage. The operators in this stage are involved in feature extraction and preliminary linear transformation of the input image, the data matrix is relatively small, while the weight matrix is large. For example, a convolution operation is performed on the input image patch to map the image features to a higher-dimensional space. For the first execution stage, the tensor parallel strategy is adopted to divide the weight matrix among multiple GPUs, and each GPU is responsible for calculating the product of a part of the weight matrix and the input data.
[0204] The operators for calculating attention scores are divided into the second execution stage. The operators in this stage are mainly used to calculate attention scores. The data matrix is relatively large because it contains the image features after preliminary transformation, while the weight matrix is relatively small. For the second execution stage, a data parallel strategy can be adopted. The input image feature data is evenly distributed to 4 GPUs, and each GPU independently calculates the attention scores for the data part it is responsible for.
[0205] During the execution of the target operator stream, in the first execution stage, each GPU performs matrix multiplication operations on the allocated part of the weight matrix and the input image data according to the tensor parallel strategy. For example, GPU1 is responsible for calculating the product of the first part of the weight matrix and the input image data, GPU2 is responsible for the second part, and so on.
[0206] After the first execution stage is completed, according to the data parallel strategy adopted in the second execution stage, the intermediate results calculated by each GPU (i.e., the image features after preliminary transformation) are divided into multiple parts by rows according to the corresponding partitioning method of the data parallel strategy, and each part is written into the memory resources of its corresponding GPU (including itself and other GPUs). Finally, each GPU has the complete image features for calculating attention scores.
[0207] Each GPU calculates the attention scores for the allocated image feature data according to the data parallel strategy. Each GPU independently completes the calculation task without excessive communication. After the second execution stage is completed, the attention score results can also be directly written into the memory resources of the corresponding GPU according to the parallel strategy of the subsequent execution stage to prepare for subsequent calculations.
[0208] In the above manner, the calculation tasks of each execution stage in the target operator stream are completed in sequence. Each stage selects an appropriate parallel strategy according to its matrix calculation scale and dependency relationship, and performs efficient data transmission between stages. After all execution stages are completed, the final calculation results of the 4 GPUs are integrated to obtain the final image recognition result. For example, by summarizing and analyzing the attention weights calculated by each GPU, it is judged whether there are specific geographical features in the image.
[0209] By dividing the target operator stream into multiple execution stages, selecting appropriate parallel strategies according to the matrix calculation scale and dependency relationship, and performing efficient data transmission between different stages, the parallel requirements of each operator can be accurately adapted, the communication time can be reduced, and thus the inference efficiency of the large model can be significantly improved. In the field of image recognition, this method can process a large amount of high-definition image data faster and meet the requirements of real-time and accuracy.
[0210] Embodiments of the present application can be implemented through a computing cluster including multiple artificial intelligence chips. Refer to Figure 4 As shown, a computing cluster including multiple GPUs is shown. Among them, the GPU contains video memory for storing data. There are multiple SPCs (Streaming Processing Clusters) in the GPU, and each SPC contains computing units. There is on-chip cache in the computing units for quickly storing and reading data to accelerate the computing process. The computing units also contain multiple execution units, and components such as physical registers are provided in the execution units, which can be used to temporarily store data in the computing process, etc.
[0211] The operator operations in each execution stage can be undertaken by the computing units of the GPU. For example, in the execution stage of calculating the attention score, the execution units in the computing units of each GPU are responsible for specific matrix multiplication and other operation operations under the data parallel strategy. The physical registers are used to temporarily store intermediate calculation data, and the on-chip cache speeds up data reading and storage, improving the operation speed.
[0212] The video memory can be used to store the input data, intermediate results and final results of each execution stage. For example, after the operation in execution stage one is completed, the intermediate result can be temporarily stored in the video memory, and then according to the data parallel strategy of the next execution stage, it is read from the video memory and directly written to the corresponding GPU. The SPC plays a role in coordinating the transfer of data between the video memory and the computing units, ensuring that the data can be accurately and timely supplied to the computing units for subsequent operator operations.
[0213] Through the collaborative work of the components in the GPU and in cooperation with the parallel strategies of different execution stages, the operation of the operator stream can be completed more efficiently, reducing the data waiting time and unnecessary communication overhead, and providing support for improving the inference efficiency of large models from the hardware level.
[0214] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0215] Based on the same inventive concept, an embodiment of the present application further provides a parallel running device for an operator flow for implementing the parallel running method of the operator flow involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the parallel running device of the operator flow provided below can refer to the limitations on the parallel running method of the operator flow in the above text, and will not be repeated here.
[0216] In an exemplary embodiment, as Figure 5 shown, a parallel running device 500 for an operator flow is provided, including: a partitioning module 502, an execution module 504, and a determination module 506, where:
[0217] The partitioning module 502 is configured to partition a target operator flow into multiple execution stages, each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy;
[0218] The execution module 504 is configured to, during the running process of the target operator flow, for any one of the execution stages, adopt the parallel strategy corresponding to the execution stage inside each artificial intelligence chip, execute the operations corresponding to the operators in the execution stage, and after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, directly write part or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip, and the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips;
[0219] The determination module 506 is configured to, after the calculations in each execution stage are completed, obtain the operation result of the target operator flow based on the final operation results of each artificial intelligence chip.
[0220] By using the parallel running device for an operator flow provided in the embodiment of the present application, during the execution process of the target operator flow, the corresponding parallel strategy can be accurately adapted according to the operator characteristics of the operators in the target operator flow, and at the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips is finely optimized, so that the operators can be fully utilized, the calculation speed of the target operator flow is significantly improved, and further the inference latency of the large model is reduced and the inference efficiency of the large model is improved.
[0221] In one of the embodiments, the conventional parallel strategy includes that each artificial intelligence chip executes the operation corresponding to the operator on the full data matrix and the full weight matrix;
[0222] The hybrid parallel strategy includes: when the data matrix is the left matrix in a multiplication operation, splitting the data matrix by columns into multiple data sub-matrices, and splitting the weight matrix by rows into multiple weight sub-matrices; or, when the data matrix is the right matrix in a multiplication operation, splitting the data matrix by rows into multiple data sub-matrices, and splitting the weights by columns into multiple weight sub-matrices; each artificial intelligence chip performs the corresponding operation of the operator based on the allocated data sub-matrix and weight sub-matrix.
[0223] In one embodiment, the execution phase includes at least one of the following:
[0224] The first execution phase, the first execution phase includes at least one first operator, the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirement, and the first execution phase corresponds to the conventional parallel strategy;
[0225] The second execution phase, the second execution phase includes at least one second operator, the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirement, but the weight matrix of the operation satisfies the parallel requirement, and the second execution phase corresponds to the tensor parallel strategy;
[0226] The third execution phase, the third execution phase includes at least one third operator, the third operator is an operator for which the data matrix of the operation satisfies the parallel requirement, but the weight matrix of the operation does not satisfy the parallel requirement, and the third execution phase corresponds to the data parallel strategy;
[0227] The fourth execution phase, the fourth execution phase includes at least one fourth operator, the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirement, and the fourth execution phase corresponds to the hybrid parallel strategy.
[0228] In one embodiment, the parallel strategy corresponding to the execution phase is the conventional parallel strategy. Based on the parallel strategy corresponding to the next execution phase, writing some or all of the operation results of each artificial intelligence chip in the execution phase directly into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0229] If the parallel strategy corresponding to the next execution phase is the tensor parallel strategy, then for any artificial intelligence chip, retain the operation result of the artificial intelligence chip in the execution phase in the memory resource of the artificial intelligence chip itself;
[0230] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with the row division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple first data sub-matrices corresponding to each artificial intelligence chip respectively, and retain the first data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself;
[0231] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively, and retain the second data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself, wherein the division method of the second data sub-matrix includes:
[0232] In the case where the data matrix in the next execution stage is the left matrix of the multiplication operation, with the column division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively; or, in the case where the data matrix in the next execution stage is the right matrix of the multiplication operation, with the row division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively.
[0233] In one embodiment, the parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0234] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0235] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, according to the number of artificial intelligence chips, with columns as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple third data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves.
[0236] In one embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0237] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0238] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fourth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0239] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fifth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0240] In one embodiment, the parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resource corresponding to the target artificial intelligence chip includes:
[0241] Perform full reduction processing on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction processing results:
[0242] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the regular parallel strategy, then for any artificial intelligence chip, retain the full reduction processing result in the memory resource of the artificial intelligence chip itself;
[0243] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, then for any artificial intelligence chip, divide the full reduction processing result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip with the row division dimension according to the number of artificial intelligence chips, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself.
[0244] Each module in the parallel operation device of the above operator flow can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0245] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for parallel operation of an operator flow. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0246] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0247] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0248] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0249] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0250] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0251] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0252] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0253] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A method for parallel operation of operator streams, characterized in that: The method comprises: Divide the target operator stream into multiple execution stages, each execution stage includes at least one operator, and each operator belonging to the same execution stage corresponds to the same parallel strategy, which includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy; During the operation of the target operator flow, for any of the execution stages, the parallel strategy corresponding to the execution stage is adopted inside each artificial intelligence chip to execute the operations corresponding to each operator in the execution stage, and after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, part or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip, and the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips; After the calculations are completed in each execution stage, the calculation results of the target operator flow are obtained based on the final calculation results of each artificial intelligence chip.
2. The method according to claim 1, characterized in that The conventional parallel strategy includes that each of the artificial intelligence chips performs operations corresponding to the operators on the entire data matrix and the entire weight matrix; The hybrid parallel strategy includes: when the data matrix is the left matrix of the multiplication operation, the data matrix is split into multiple data sub-matrices with columns as the splitting dimension, and the weight matrix is split into multiple weight sub-matrices with rows as the splitting dimension; or, when the data matrix is the right matrix of the multiplication operation, the data matrix is split into multiple data sub-matrices with rows as the splitting dimension, and the weight is split into multiple weight sub-matrices with columns as the splitting dimension; each artificial intelligence chip executes the operation corresponding to the operator according to the data sub-matrix and the weight sub-matrix allocated to it.
3. The method according to claim 2, characterized in that The execution phase includes at least one of the following: A first execution stage, the first execution stage includes at least one first operator, the first operator is an operator whose data matrix and weight matrix do not meet the parallel requirements, and the first execution stage corresponds to the conventional parallel strategy; A second execution stage, the second execution stage includes at least one second operator, the second operator is an operator whose data matrix does not meet the parallel requirement but whose weight matrix meets the parallel requirement, and the second execution stage corresponds to the tensor parallel strategy; A third execution stage, the third execution stage includes at least one third operator, the third operator is an operator whose data matrix of the operation meets the parallel requirement, but whose weight matrix of the operation does not meet the parallel requirement, and the third execution stage corresponds to the data parallel strategy; The fourth execution stage includes at least one fourth operator, and the fourth operator is an operator whose data matrix and weight matrix both meet the parallel requirements. The fourth execution stage corresponds to the hybrid parallel strategy.
4. The method according to claim 2 or 3, characterized in that: The parallel strategy corresponding to the execution stage is the conventional parallel strategy, and the parallel strategy corresponding to the next execution stage is based on writing part or all of the calculation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip, including at least one of the following: If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy, then for any of the artificial intelligence chips, the calculation results of the artificial intelligence chip in the execution stage are retained in the memory resources of the artificial intelligence chip itself; If the parallel strategy corresponding to the next execution stage is the data parallel strategy, then for any of the artificial intelligence chips, based on the number of artificial intelligence chips and the behavior division dimension, the calculation results of the artificial intelligence chip in the execution stage are divided into a plurality of first data sub-matrices corresponding to each artificial intelligence chip, and the first data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resource of the artificial intelligence chip itself; If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, then for any of the artificial intelligence chips, according to the number of artificial intelligence chips, the calculation results of the artificial intelligence chip in the execution stage are divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively, and the second data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resource of the artificial intelligence chip itself, wherein the division method of the second data sub-matrix includes: When the data matrix of the next execution stage is a left matrix of multiplication operation, the columns are used as the partitioning dimension to divide the operation results of the artificial intelligence chip in the execution stage into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively; or, when the data matrix of the next execution stage is a right matrix of multiplication operation, the rows are used as the partitioning dimension to divide the operation results of the artificial intelligence chip in the execution stage into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively.
5. The method according to claim 2 or 3, characterized in that: The parallel strategy corresponding to the execution stage is the data parallel strategy, and the parallel strategy corresponding to the next execution stage is based on writing part or all of the calculation results of each artificial intelligence chip in the execution stage directly into the memory resources of the target artificial intelligence chip, including at least one of the following: If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, then for any of the artificial intelligence chips, the calculation results of the artificial intelligence chip in the execution stage are written into the memory resources of each artificial intelligence chip respectively; If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, then for any of the artificial intelligence chips, when the data matrix of the next execution stage is a left matrix of multiplication operation, the calculation results of the artificial intelligence chip in the execution stage are divided into multiple third data sub-matrices corresponding to each artificial intelligence chip according to the number of artificial intelligence chips and with columns as the division dimension, and the multiple third data sub-matrices are written into the memory resources of the corresponding artificial intelligence chips respectively; or, when the data matrix of the next execution stage is a right matrix of multiplication operation, the calculation results of the artificial intelligence chip in the execution stage are retained in the memory resources of the artificial intelligence chip itself.
6. The method according to claim 2 or 3, characterized in that: The parallel strategy corresponding to the execution stage is the tensor parallel strategy, and the parallel strategy corresponding to the next execution stage is based on writing part or all of the calculation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip, including at least one of the following: If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, then for any of the artificial intelligence chips, the calculation results of the artificial intelligence chip in the execution stage are written into the memory resources of each artificial intelligence chip respectively; If the parallel strategy corresponding to the next execution stage is the data parallel strategy, then for any of the artificial intelligence chips, based on the number of artificial intelligence chips and the behavior division dimension, the calculation results of the artificial intelligence chip in the execution stage are divided into a plurality of fourth data sub-matrices corresponding to each artificial intelligence chip, and the plurality of fourth data sub-matrices are respectively written into the memory resources of the artificial intelligence chips corresponding to each of the artificial intelligence chips; If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, then for any of the artificial intelligence chips, when the data matrix of the next execution stage is the left matrix of the multiplication operation, the operation results of the artificial intelligence chip in the execution stage are retained in the memory resources of the artificial intelligence chip itself; or, when the data matrix of the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, the operation results of the artificial intelligence chip in the execution stage are divided into multiple fifth data sub-matrices corresponding to each artificial intelligence chip respectively according to the behavior division dimension, and the multiple fifth data sub-matrices are respectively written into the memory resources of the corresponding artificial intelligence chips.
7. The method according to claim 2 or 3, characterized in that: The parallel strategy corresponding to the execution stage is the hybrid parallel strategy, and the parallel strategy corresponding to the next execution stage is based on writing part or all of the calculation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip, including: Performing full reduction processing on the calculation results of each of the artificial intelligence chips in the execution phase, and performing at least one of the following on the full reduction processing results: If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, then for any of the artificial intelligence chips, the full reduction processing result is retained in the memory resources of the artificial intelligence chip itself; If the parallel strategy corresponding to the next execution stage is the data parallel strategy, then for any of the artificial intelligence chips, based on the number of artificial intelligence chips and the behavior division dimension, the full reduction processing result is divided into multiple sixth data sub-matrices corresponding to each artificial intelligence chip, and the sixth data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resources of the artificial intelligence chip itself.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Parallel strategy search method for efficient training of artificial intelligence large model
CN116680301A
Model reasoning method and device, equipment and storage medium
CN118133964A
Model operator parallel splitting method and device, equipment and storage medium
CN119294463A
Conditional parallel processing in fully-connected neural networks
US20170193368A1
Method and apparatus for accelerating convolutional neural network
US20230289230A1