Parallel operation method of operator flow, computer device and readable storage medium
By dividing the operator flow in the large model inference process into multiple execution stages, parallel strategy and result delivery optimization, the problem of unbalanced computing power utilization is solved, and more efficient computing speed and reduced latency are achieved.
Patent Information
- Application Number
- CN202510653876.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the process of large-scale model inference, a single data parallel or tensor parallel strategy in the prior art leads to uneven computing power utilization, resulting in idle hardware resources and difficult to meet the strict requirements of low latency.
The target operator stream is divided into multiple execution stages. Each stage adopts a suitable parallel strategy (regular, data, tensor, mixed parallel strategy). After the operation is completed in the execution stage, the operation results are directly written into the memory resources of the artificial intelligence chip according to the parallel strategy of the next stage to optimize the way the operation results are transferred between chips.
By accurately adapting parallel strategies and optimizing the transmission method of operation results, the operator resources are fully utilized, the computing speed is significantly improved, the inference delay of large models is reduced, and the inference efficiency is improved.
Smart Images

Figure CN120179295B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence chips, and particularly to a method for parallel running of operator streams, a computer device, and a readable storage medium. Background Technique
[0002] In the scenario of large model inference, efficiently utilizing hardware resources to reduce inference latency is the key to improving model performance. Data parallelism and tensor parallelism are two common strategies for accelerating large model calculations. Among them, data parallelism divides the input data into different computing devices (such as GPUs (Graphics Processing Units)) for parallel processing, and tensor parallelism divides the model weights among the computing devices for collaborative calculation. Both aim to give full play to the computing power of operators.
[0003] However, the large model inference process involves multiple computing links, and there may be significant differences in the matrix calculation scales and characteristics involved in each link. In response to this situation, if a single data parallelism or tensor parallelism strategy is adopted, it often leads to unbalanced computing power utilization. For example, for the computing tasks corresponding to some operators, the hardware resources can be fully utilized; but for the computing tasks corresponding to other operators, due to the data partitioning or tensor segmentation method not matching the computing tasks of the operator, a large amount of computing power will be idle, and the efficient utilization of computing resources cannot be achieved, making it difficult to meet the stringent requirements of large model inference for low latency. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method for parallel running of operator streams, a computer device, and a readable storage medium that can reduce the inference latency of large models.
[0005] In a first aspect, this application provides a method for parallel running of operator streams, including:
[0006] Dividing a target operator stream into multiple execution stages, each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy;
[0007] During the running of the target operator stream, for any one of the execution stages, adopt the parallel strategy corresponding to the execution stage inside each artificial intelligence chip to execute the operations corresponding to the operators in the execution stage, and after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, directly write part or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip, and the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips;
[0008] After calculations are completed in each execution stage, the operation result of the target operator stream is obtained based on the final operation results of the artificial intelligence chips.
[0009] In one embodiment, the conventional parallel strategy includes that each artificial intelligence chip performs the operation corresponding to the operator on the full data matrix and the full weight matrix;
[0010] The hybrid parallel strategy includes: when the data matrix is the left matrix of the multiplication operation, the data matrix is split into multiple data sub-matrices with columns as the splitting dimension, and the weight matrix is split into multiple weight sub-matrices with rows as the splitting dimension; or, when the data matrix is the right matrix of the multiplication operation, the data matrix is split into multiple data sub-matrices with rows as the splitting dimension, and the weights are split into multiple weight sub-matrices with columns as the splitting dimension; each artificial intelligence chip performs the operation corresponding to the operator based on the allocated data sub-matrix and the weight sub-matrix.
[0011] In one embodiment, the execution stage includes at least one of the following:
[0012] The first execution stage, the first execution stage includes at least one first operator, the first operator is an operator for which neither the data matrix nor the weight matrix of the operation meets the parallel requirements, and the first execution stage corresponds to the conventional parallel strategy;
[0013] The second execution stage, the second execution stage includes at least one second operator, the second operator is an operator for which the data matrix of the operation does not meet the parallel requirements, but the weight matrix of the operation meets the parallel requirements, and the second execution stage corresponds to the tensor parallel strategy;
[0014] The third execution stage, the third execution stage includes at least one third operator, the third operator is an operator for which the data matrix of the operation meets the parallel requirements, but the weight matrix of the operation does not meet the parallel requirements, and the third execution stage corresponds to the data parallel strategy;
[0015] The fourth execution stage, the fourth execution stage includes at least one fourth operator, the fourth operator is an operator for which both the data matrix and the weight matrix of the operation meet the parallel requirements, and the fourth execution stage corresponds to the hybrid parallel strategy.
[0016] In one embodiment, the parallel strategy corresponding to the execution stage is the conventional parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0017] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy, for any one of the artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is stored in the memory resource of the artificial intelligence chip itself;
[0018] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, taking the row as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of first data sub-matrices corresponding to each artificial intelligence chip respectively, and the first data sub-matrix corresponding to the artificial intelligence chip itself is stored in the memory resource of the artificial intelligence chip itself;
[0019] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively, and the second data sub-matrix corresponding to the artificial intelligence chip itself is stored in the memory resource of the artificial intelligence chip itself, wherein the division method of the second data sub-matrix includes:
[0020] When the data matrix in the next execution stage is the left matrix of the multiplication operation, taking the column as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, taking the row as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively.
[0021] In one embodiment, the parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0022] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is written into the memory resources of each artificial intelligence chip respectively;
[0023] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, according to the number of artificial intelligence chips, with columns as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple third data sub-matrices corresponding to each artificial intelligence chip respectively, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves.
[0024] In one embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0025] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0026] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fourth data sub-matrices corresponding to each artificial intelligence chip respectively, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0027] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fifth data sub-matrices corresponding to each artificial intelligence chip respectively, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0028] In one embodiment, the parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes:
[0029] Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction process results:
[0030] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, then for any artificial intelligence chip, retain the full reduction process results in the memory resources of the artificial intelligence chip itself;
[0031] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, then for any artificial intelligence chip, divide the full reduction process results into multiple sixth data sub-matrices corresponding to each artificial intelligence chip in the dimension of row division according to the number of artificial intelligence chips, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself.
[0032] In a second aspect, the present application also provides a parallel operation device for an operator stream, including:
[0033] A division module, configured to divide a target operator stream into multiple execution stages, each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy;
[0034] An execution module, configured to, during the operation of the target operator stream, for any execution stage, adopt the parallel strategy corresponding to the execution stage inside each artificial intelligence chip to execute the operations corresponding to the operators in the execution stage, and after the execution stage operation is completed, based on the parallel strategy corresponding to the next execution stage, write some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip, and the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips;
[0035] A determination module, configured to obtain the operation result of the target operator stream based on the final operation results of each artificial intelligence chip after the calculations of all execution stages are completed.
[0036] In one embodiment, the conventional parallel strategy includes each of the artificial intelligence chips performing the operations corresponding to the operators on the full data matrix and the full weight matrix;
[0037] The hybrid parallel strategy: when the data matrix is the left matrix in a multiplication operation, the data matrix is split into multiple data sub - matrices with columns as the splitting dimension, and the weight matrix is split into multiple weight sub - matrices with rows as the splitting dimension; or, when the data matrix is the right matrix in a multiplication operation, the data matrix is split into multiple data sub - matrices with rows as the splitting dimension, and the weights are split into multiple weight sub - matrices with columns as the splitting dimension; each artificial intelligence chip performs the operations corresponding to the operators based on the allocated data sub - matrices and weight sub - matrices.
[0038] In one embodiment, the execution phase includes at least one of the following:
[0039] The first execution phase, the first execution phase includes at least one first operator, the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirements, and the first execution phase corresponds to the conventional parallel strategy;
[0040] The second execution phase, the second execution phase includes at least one second operator, the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirements, but the weight matrix of the operation satisfies the parallel requirements, and the second execution phase corresponds to the tensor parallel strategy;
[0041] The third execution phase, the third execution phase includes at least one third operator, the third operator is an operator for which the data matrix of the operation satisfies the parallel requirements, but the weight matrix of the operation does not satisfy the parallel requirements, and the third execution phase corresponds to the data parallel strategy;
[0042] The fourth execution phase, the fourth execution phase includes at least one fourth operator, the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirements, and the fourth execution phase corresponds to the hybrid parallel strategy.
[0043] In one embodiment, when the parallel strategy corresponding to the execution phase is the conventional parallel strategy, and based on the parallel strategy corresponding to the next execution phase, writing some or all of the operation results of each artificial intelligence chip in the execution phase directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0044] If the parallel strategy corresponding to the next execution phase is the tensor parallel strategy, then for any artificial intelligence chip, the operation result of the artificial intelligence chip in the execution phase is retained in the memory resources of the artificial intelligence chip itself;
[0045] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with the row division dimension, divide the operation results of the artificial intelligence chips in the execution stage into a plurality of first data sub-matrices respectively corresponding to each artificial intelligence chip, and retain the first data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself;
[0046] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, divide the operation results of the artificial intelligence chips in the execution stage into a plurality of second data sub-matrices respectively corresponding to each artificial intelligence chip, and retain the second data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself, wherein the division method of the second data sub-matrix includes:
[0047] In the case where the data matrix in the next execution stage is the left matrix of the multiplication operation, with the column division dimension, divide the operation results of the artificial intelligence chips in the execution stage into a plurality of second data sub-matrices respectively corresponding to each artificial intelligence chip; or, in the case where the data matrix in the next execution stage is the right matrix of the multiplication operation, with the row division dimension, divide the operation results of the artificial intelligence chips in the execution stage into a plurality of second data sub-matrices respectively corresponding to each artificial intelligence chip.
[0048] In one embodiment, the parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing part or all of the operation results of each artificial intelligence chip in the execution stage into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0049] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0050] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, then for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, according to the number of artificial intelligence chips, with columns as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple third data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves.
[0051] In one embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0052] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, then for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0053] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, then for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fourth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0054] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, then for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fifth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0055] In one embodiment, the parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resource corresponding to the target artificial intelligence chip includes:
[0056] Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction process results:
[0057] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, for any artificial intelligence chip, retain the full reduction process results in the memory resource of the artificial intelligence chip itself;
[0058] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any artificial intelligence chip, divide the full reduction process results into multiple sixth data sub-matrices corresponding to each artificial intelligence chip in the dimension of row division according to the number of artificial intelligence chips, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip.
[0059] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the parallel operation method of the operator stream in any one of the above.
[0060] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the parallel operation method of the operator stream in any one of the above.
[0061] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the parallel operation method of the operator stream in any one of the above.
[0062] The above-mentioned parallel operation method, computer device, and readable storage medium of the operator flow can divide the target operator flow into multiple execution stages. Each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy. The parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy. Furthermore, during the operation of the target operator flow, for any execution stage, the parallel strategy corresponding to the execution stage is adopted inside each artificial intelligence chip to execute the operations corresponding to the operators within the execution stage. And after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, part or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip. The target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips. After the calculations are completed in each execution stage, the operation result of the target operator flow is obtained based on the final operation results of each artificial intelligence chip. By using the parallel operation method, computer device, and readable storage medium of the operator flow provided in the embodiments of the present application, during the execution process of the target operator flow, the corresponding parallel strategy can be accurately adapted according to the operator characteristics of the operators in the target operator flow. At the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips is finely optimized, so that the operators can be fully utilized, the calculation speed of the target operator flow can be significantly improved, and thus the inference latency of the large model can be reduced and the inference efficiency of the large model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0064] Figure 1 It is a schematic diagram of the attention operator flow in an embodiment;
[0065] Figure 2 It is a schematic flowchart of the parallel operation method of the operator flow in an embodiment;
[0066] Figure 3 It is a schematic diagram of the parallel processing of the attention operator flow in an embodiment;
[0067] Figure 4 It is a schematic diagram of the computing cluster in another embodiment;
[0068] Figure 5 It is a structural block diagram of the parallel operation device of the operator flow in an embodiment;
[0069] Figure 6 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0070] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0071] The large model inference process includes multiple computing links, and there may be obvious differences in the matrix computing scale and characteristics involved in each link. Taking the attention operator flow (i.e., the attention operator flow) in the large model inference process as an example, the embodiments of the present application will be described below. It should be understood that the attention operator flow is only an example of the target operator flow in the embodiments of the present application, and is not understood as a limitation on the target operator flow. In fact, any operator flow with matrix scale and characteristic transformation in the computing link is applicable to the embodiments of the present application.
[0072] The attention operator flow is composed of a first MMA (Matrix Multiply - Accumulate) operator, a split - root mean square normalization operator (which can also be marked as the split_rmsnorm operator, an operator that combines the operations of splitting and root mean square normalization), a second MMA operator, a rotary position embedding operator (which can also be marked as ROPE (Rotary Position Embedding)), a third MMA operator, a normalized exponential function operator (i.e., the softmax operator), a fourth MMA operator, and an attention multiplication matrix (which can also be marked as attn_MMA) operator.
[0073] For the attention operator flow, if a single parallel strategy is adopted, it will lead to the problem that some operators cannot be fully utilized, resulting in waste of operator resources. Taking the tensor parallel strategy as an example, referring to Figure 1 as shown, it shows a schematic diagram of an attention operator flow performing operations using the tensor parallel strategy.
[0074] It should be noted that the attention operator and the shape expressions of different tensors in the embodiments of this application are only used as an example implementation in the embodiments of this application, and should not be understood as a limitation on the attention operator and the shape expressions of tensors. In fact, for attention operators with different structures, there is a corresponding relationship between the shape expressions of their tensors and the parameters of the attention operator, which is not specifically limited in the embodiments of this application.
[0075] Refer to Figure 1 the example shown below. First, in the input stage, the scale of the input data is (B, 7168), where B represents the batch size. Enter the first MMA operator. In the current operation stage, the scale of the weight matrix is [7168, 1536 + 576]. Perform matrix multiplication on the input data to obtain an output tensor qc with a shape of (B, 1536 + 576).
[0076] In the normalization and splitting stage, the output tensor qc enters the split-root mean square normalization operator, which splits the output tensor qc into two parts. One part is a tensor q_c with a shape of (B, 1536), and the other part is a tensor kv_c with a shape of (B, 576).
[0077] In the feature processing stage, in this example, it is assumed that there are 4 GPUs (Graphics Processing Units) in the cluster. The tensor q_c enters the second MMA operator. Since the current tensor parallel strategy is adopted, in the current operation stage, the scale of the weight matrix is [1536, 128×576]. Then, the part allocated to each GPU is a weight matrix with a scale of [1536, 32×576]. Perform matrix multiplication based on this weight matrix, and then pass through the rotary position embedding operator to perform rotary encoding processing only on part of the data, obtaining a tensor q with a shape of (B×32, 1, 576). It should be noted that Figure 1 for the 4 GPUs in this operation stage, only one of them is used as an example to show the operation part it executes, and the other 3 GPUs are not shown. Refer to the operation part executed by the shown GPU.
[0078] The tensor kv_c enters the rotary position embedding module to perform rotary encoding processing only on part of the data in 576 dimensions, obtaining a tensor kv with a shape of (B, S, 576), where S is the sequence length.
[0079] In the attention calculation stage, tensor q and tensor kv enter the third MMA operator together for matrix multiplication, obtaining a tensor s1 with a shape of (B×32, 1, S). Tensor s1 enters the normalization exponential function operator for normalization operation, and the output tensor s2 still has a shape of (B×32, 1, S), which is used to calculate the attention weights. After the normalization operation, s2 and tensor kv enter the fourth MMA operator together for matrix multiplication operation, obtaining a tensor O with a shape of (B, 32×512).
[0080] Finally, tensor O enters the attention multiplication matrix operator. The weight matrix shape in the current operation stage is [32×512, 7168]. After matrix multiplication operation, the final output data with a shape of (B, 7168) is obtained, completing the entire calculation process.
[0081] It can be seen that the tensor parallel strategy can make full use of the second MMA operator and the attention multiplication matrix operator. However, for other computing powers, such as the third MMA operator and the fourth MMA operator, they are not fully utilized, resulting in idle computing power. This situation may affect the low-latency performance of large model inference and make it difficult to meet strict latency requirements.
[0082] The embodiment of the present application provides a method for parallel running of an operator stream. In the execution process of the target operator stream, by dividing the target operator stream into multiple execution stages according to the operator characteristics of each operator in the target operator stream, and precisely matching a corresponding parallel strategy for each execution stage. At the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips can be finely optimized, so as to make full use of each operator in the target operator stream, significantly improve the calculation speed of the target operator stream, and then reduce the large model inference latency and improve the inference efficiency of the large model.
[0083] In an exemplary embodiment, as Figure 2 shown, a method for parallel running of an operator stream is provided. Taking the host side as an example for illustration, it can be understood that the host side may include a CPU (Central Processing Unit, central processing unit). The method includes the following steps 202 to 206. Wherein:
[0084] Step 202, divide the target operator stream into multiple execution stages. Each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy. The parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy.
[0085] In the embodiments of the present application, the target operator stream can be divided into multiple execution stages based on the matrix calculation scale and characteristics of different operators in the target operator stream, as well as the dependency relationships between the operators. Among them, the principle of division is to make the operators within each execution stage have similarity in the requirements for computing resources and computing characteristics, so as to adopt the same parallel strategy for optimization.
[0086] Exemplarily, starting from the starting operator of the target operator stream, it can be used as the starting operator of the first execution stage. Further, according to the dependency relationships between the operators, the subsequent operators are sequentially added to the current execution stage until an operator is encountered whose matrix calculation scale or calculation characteristics are significantly different from those of the operators within the current stage, or there is a complex dependency relationship between this operator and the operators within the current stage, and the same parallel strategy cannot be adopted. At this time, this operator is used as the starting operator of the next execution stage, and the above process is repeated until all the operators in the target operator stream are divided into the corresponding execution stages.
[0087] In an exemplary embodiment, the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy, where: the conventional parallel strategy includes each artificial intelligence chip performing the operation corresponding to the operator on the full data matrix and the full weight matrix.
[0088] The hybrid parallel strategy includes, when the data matrix is the left matrix in a multiplication operation, splitting the data matrix into multiple data sub-matrices with columns as the splitting dimension and splitting the weight matrix into multiple weight sub-matrices with rows as the splitting dimension; or, when the data matrix is the right matrix in a multiplication operation, splitting the data matrix into multiple data sub-matrices with rows as the splitting dimension and splitting the weight matrix into multiple weight sub-matrices with columns as the splitting dimension; each artificial intelligence chip performs the operation corresponding to the operator based on the allocated data sub-matrix and weight sub-matrix.
[0089] In the embodiments of the present application, the artificial intelligence chip includes chips such as GPU (Graphics Processing Unit), GPGPU (General-Purpose Computing on Graphics Processing Units), TPU (Tensor Processing Unit), and NPU (Neural network Processing Unit). In the following embodiments of the present application, the artificial intelligence chip will be taken as an example of GPU for illustration. Taking a cluster composed of 4 GPUs as an example, the conventional parallel strategy, data parallel strategy, tensor parallel strategy, and hybrid parallel strategy are described as follows:
[0090] Conventional parallel strategy: Without any partitioning of the data matrix, the complete data matrix is directly transmitted to 4 GPUs. Each GPU uses the same unpartitioned weight matrix and performs the same arithmetic operation on the received data matrix. Finally, the 4 GPUs output the same arithmetic result.
[0091] Data Parallelism (DP) strategy: The data matrix is partitioned. The complete data matrix is split into 4 parts with approximately the same amount of data in each part, and then these 4 parts of the data matrix are written into 4 GPUs respectively. The weight matrix is not partitioned, and the same weight matrix is stored in 4 GPUs respectively. Each GPU performs arithmetic operations on the received partial data matrix and the same weight matrix to obtain its respective partial arithmetic result. Finally, after these partial arithmetic results are aggregated and combined, the final complete arithmetic result is obtained.
[0092] Tensor Parallelism (TP) strategy: The weight matrix is partitioned. For example, it is partitioned into 4 parts by columns with approximately the same amount of data in each part, and then these 4 parts of the weight matrix are written into 4 GPUs respectively. The data matrix is not partitioned, and the complete data matrix is written into 4 GPUs respectively. Each GPU performs arithmetic operations on the complete data matrix using the partial weight matrix it stores. Finally, after the arithmetic results of each GPU are aggregated and combined, the complete arithmetic result can be obtained.
[0093] Hybrid parallel strategy: When the data matrix is the left matrix in a multiplication operation, that is, when the data matrix is the left matrix and the weight matrix is the right matrix, the data matrix is split into 4 data sub - matrices with approximately the same amount of data by taking columns as the splitting dimension; at the same time, the weight matrix is split into 4 weight sub - matrices by taking rows as the splitting dimension. Subsequently, these 4 data sub - matrices and 4 weight sub - matrices are respectively assigned to 4 GPUs. Or, when the data matrix is the right matrix in a multiplication operation, that is, when the data matrix is the right matrix and the weight matrix is the left matrix, the data matrix is split into 4 data sub - matrices by taking rows as the splitting dimension, and the weight is split into 4 weight sub - matrices by taking columns as the splitting dimension. Each GPU performs arithmetic operations on the received data sub - matrix and weight sub - matrix to complete the calculation of its respective responsible part. Finally, the arithmetic results of the 4 GPUs are aggregated and integrated to obtain the final complete arithmetic result.
[0094] It should be noted that the splitting process of the above data matrix and weight matrix can be an equal split, that is, the sizes of the split sub-matrices are exactly the same, or it can be split based on the performance resources of the GPU, that is, the sizes of the split sub-matrices are not exactly the same and are adapted to the performance resources of their respective GPUs. That is, a relatively larger sub-matrix can be divided for a GPU with better performance resources, and a relatively smaller sub-matrix can be divided for a GPU with poorer performance resources, so as to make full use of GPU resources.
[0095] In the embodiments of the present application, after dividing the target operator stream into multiple execution stages, a corresponding parallel strategy can be matched for each execution stage. In an exemplary embodiment, the execution stage may include at least one of a first execution stage, a second execution stage, a third execution stage, and a fourth execution stage, where:
[0096] The first execution stage includes at least one first operator, and the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirement. The first execution stage corresponds to a conventional parallel strategy;
[0097] The second execution stage includes at least one second operator, and the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirement, but the weight matrix of the operation satisfies the parallel requirement. The second execution stage corresponds to a tensor parallel strategy;
[0098] The third execution stage includes at least one third operator, and the third operator is an operator for which the data matrix of the operation satisfies the parallel requirement, but the weight matrix of the operation does not satisfy the parallel requirement. The third execution stage corresponds to a data parallel strategy;
[0099] The fourth execution stage includes at least one fourth operator, and the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirement. The fourth execution stage corresponds to a hybrid parallel strategy.
[0100] In the embodiments of the present application, in the first execution stage, there is at least one first operator. The data matrix and weight matrix of such first operators are both relatively small. Due to the limited matrix size, it is difficult to significantly improve the efficiency through parallel computing, that is, neither of them satisfies the parallel computing conditions. Based on this characteristic, it can be determined that the first execution stage adopts a conventional parallel strategy. Under this conventional parallel strategy, the complete data matrix and weight matrix will be written to each GPU respectively, and each GPU will execute the same full-scale operation, and finally output the same calculation result.
[0101] In the second execution stage, there is at least one second operator. Such second operators exhibit the characteristics of a relatively small data matrix size and a relatively large weight matrix size. Specifically, due to the small size of its data matrix, it cannot fully utilize parallel computing resources to improve efficiency and does not meet the requirements of parallel computing; while the weight matrix, because of its large size, has the conditions for parallel computing. Based on this characteristic, it can be determined that the second execution stage adopts a tensor parallel strategy. Under this tensor parallel strategy, each GPU divides the weight matrix and stores it distributively, and at the same time each GPU receives the complete data matrix. Each GPU uses the stored partial weight matrix and the complete data matrix for operations, and can efficiently complete the operation processing tasks of the second operator.
[0102] In the third execution stage, there is at least one third operator. Such third operators exhibit the characteristics of a relatively large data matrix size and a relatively small weight matrix size. Specifically, due to the small size of its weight matrix, it cannot fully utilize parallel computing resources to improve efficiency and does not meet the requirements of parallel computing; while the data matrix, because of its large size, has the conditions for parallel computing. Based on this characteristic, it can be determined that the third execution stage adopts a data parallel strategy. Under this data parallel strategy, the data matrix is split into multiple parts and distributed to each GPU respectively; while the complete weight matrix is copied to each GPU. Each GPU uses the assigned partial data matrix and the complete weight matrix for operations, thereby efficiently completing the operation processing of the third operator.
[0103] In the fourth execution stage, there is at least one fourth operator. Such fourth operators exhibit the characteristics of both a relatively large data matrix size and a relatively large weight matrix size. Specifically, due to their large sizes, both the weight matrix and the data matrix have the conditions for parallel computing. Based on this characteristic, it can be determined that the fourth execution stage adopts a hybrid parallel strategy. Under this hybrid parallel strategy, each GPU divides the weight matrix and stores it distributively, and the data matrix is split into multiple parts and distributed to each GPU respectively. Each GPU uses the assigned partial data matrix and partial weight matrix for operations, thereby efficiently completing the operation processing of the fourth operator.
[0104] Step 204, during the running of the target operator stream, for any execution stage, within each artificial intelligence chip, adopt the parallel strategy corresponding to the execution stage to execute the operations corresponding to each operator within the execution stage. And after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, directly write some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip, where the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips.
[0105] In the embodiments of the present application, when the target operator stream is in the running (i.e., computing) state, starting from the starting operator, operations are carried out according to the parallel strategy corresponding to its execution stage. After the operations in the current execution stage are completed, based on the parallel strategy of the next execution stage, part or all of the content of the GPU operation results in the current execution stage is directly written to the memory resources of the target GPU. Here, the direct write operation means that only the data required for the target GPU to operate in the next execution stage is transmitted through the communication link between GPUs. The embodiments of the present application do not specifically limit the communication method between GPUs. For example, if there is a P2P link (Peer-to-Peer Link) interconnection between GPUs, communication can be carried out through this P2P link interconnection; if there is no P2P link interconnection, data communication can also be achieved using the PCIE (Peripheral Component Interconnect Express, high-speed peripheral component interconnection standard) bus.
[0106] In this way, the amount of data transferred between GPUs can be effectively reduced, the communication time consumption can be reduced, and thus the inference latency can be reduced, significantly improving the inference efficiency of the large model. The specific process of the direct write operation will be elaborated in detail in the following embodiments, and the embodiments of the present application will not describe it in detail here.
[0107] Step 206, after the calculations in each execution stage are completed, the operation result of the target operator stream is obtained based on the final operation results of each artificial intelligence chip.
[0108] In the embodiments of the present application, after the calculations in each execution stage are completed, that is, after each GPU has completed the operation of the target operator stream and obtained the final operation result, the operation result of the target operator stream can be obtained by summarizing and merging the final operation results of each GPU. The specific summarizing and merging method can be determined based on the parallel strategy adopted in the last stage of each GPU. The embodiments of the present application will not elaborate on the summarizing and merging method.
[0109] By adopting the parallel operation method of the operator stream provided by the embodiments of the present application, during the execution process of the target operator stream, the corresponding parallel strategy can be accurately adapted according to the operator characteristics of each operator in the target operator stream, and at the same time, based on the parallel strategies adopted in different execution stages, the direct write method of the operation results between different artificial intelligence chips can be finely optimized, so that the operators can be fully utilized, significantly improving the calculation speed of the target operator stream, thereby reducing the inference latency of the large model and improving the inference efficiency of the large model.
[0110] In an exemplary embodiment, the parallel strategy corresponding to the execution stage is a conventional parallel strategy. Based on the parallel strategy corresponding to the next execution stage, some or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resource corresponding to the target artificial intelligence chip, including at least one of the following:
[0111] If the parallel strategy corresponding to the next execution stage is a tensor parallel strategy, for any artificial intelligence chip, the operation result of the artificial intelligence chip in the execution stage is retained in the memory resource of the artificial intelligence chip itself;
[0112] If the parallel strategy corresponding to the next execution stage is a data parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, with the row as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into multiple first data sub-matrices corresponding to each artificial intelligence chip respectively, and the first data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resource of the artificial intelligence chip itself;
[0113] If the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, the operation result of the artificial intelligence chip in the execution stage is divided into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively, and the second data sub-matrix corresponding to the artificial intelligence chip itself is retained in the memory resource of the artificial intelligence chip itself, where the division method of the second data sub-matrix includes:
[0114] When the data matrix in the next execution stage is the left matrix of the multiplication operation, with the column as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, with the row as the division dimension, the operation result of the artificial intelligence chip in the execution stage is divided into multiple second data sub-matrices corresponding to each artificial intelligence chip respectively.
[0115] In the embodiment of the present application, it is assumed that the current execution stage corresponds to the first execution stage, and its corresponding parallel strategy is a conventional parallel strategy. The conventional parallel strategy refers to the operation between the full data matrix A (scale (m, k)) and the full weight matrix B (scale (k, n)). After the operation is completed, each GPU will obtain the same operation result, and this operation result is a complete data matrix C (scale (m, n)), which is the input data for the next execution stage.
[0116] In one example, the parallel strategy corresponding to the next execution stage is the tensor parallel strategy. Since the tensor parallel strategy performs operations using the complete data matrix and a partial weight matrix stored inside the GPU, at this time, each GPU can calculate the complete data matrix C without data synchronization with other GPUs. Therefore, each GPU can directly write the operation result into its own memory resource at this time.
[0117] In another example, the parallel strategy corresponding to the next execution stage is the data parallel strategy. The data parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs, and then writes these multiple data matrices into each GPU respectively for operations with the full amount of weight matrices. Since each GPU calculates the complete data matrix at this time, the operation result can be directly divided into multiple first data sub-matrices with the row as the division dimension. Each first data sub-matrix corresponds to one GPU, and the first data sub-matrix corresponding to each GPU itself is retained in the memory resource of the GPU itself.
[0118] Exemplarily, assume that there are 4 GPUs in the cluster, namely GPU0, GPU1, GPU2, and GPU3. The operation result obtained by each GPU in the current execution stage is a data matrix C1 with a size of (m, n). If the data parallel strategy is adopted in the next execution stage, each GPU should operate on a partial data matrix with a size of (m / 4, n). Then, for the operation result data matrix C1 of each GPU, a division operation with the row as the division dimension can be performed to obtain 4 first data sub-matrices with a size of (m / 4, n):
[0119] First data sub-matrix C 10 : corresponding to the 1st to m / 4th rows of C1;
[0120] First data sub-matrix C 11 : corresponding to the (m / 4 + 1)th to m / 2th rows of C1;
[0121] First data sub-matrix C 12 : corresponding to the (m / 2 + 1)th to 3m / 4th rows of C1;
[0122] First data sub-matrix C 13 : corresponding to the (3m / 4 + 1)th to mth rows of C1.
[0123] Furthermore, C 10 can be retained in the memory resource of GPU0 itself, C 11 can be retained in the memory resource of GPU1 itself, C 12 can be retained in the memory resource of GPU2 itself, C 13It is stored in the memory resources of the GPU 3 itself. In the computing process of the next execution stage, the first data sub-matrix written and the full weight matrix B2 (with a size of (n, p)) are respectively computed within each GPU to obtain a computing result C with a size of (m / 4, p). 20 and C 21 and C 22 and C 23 . Among them, C 20 corresponds to the 1st to m / 4th rows of the final computing result C2, C 21 corresponds to the (m / 4 + 1)th to m / 2th rows of the final computing result C2, C 22 corresponds to the (m / 2 + 1)th to 3m / 4th rows of the final computing result C2, C 23 corresponds to the (3m / 4 + 1)th to mth rows of the final computing result C2. C 20 and C 21 and C 22 and C 23 After summarizing and merging, it is the final computing result C2 (with a size of (m, p)) of the operator in the next execution stage.
[0124] In another example, the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy. The hybrid parallel strategy divides the complete data matrix into multiple data matrices by column or row as the division dimension according to the number of GPUs, and divides the complete weight matrix into multiple weight matrices by row or column as the division dimension and stores them in each GPU. Then, the multiple data matrices are respectively written into each GPU to compute with the partial weight matrix stored in it. Since the GPUs all compute to obtain the complete data matrix at this time, the operation results can be directly divided into multiple second data sub-matrices by column or row as the division dimension. Each second data sub-matrix corresponds to a GPU, and the second data sub-matrix corresponding to each GPU itself is stored in the memory resources of the GPU itself.
[0125] In one example, in the multiplication operation process of the next execution stage, the data matrix is the left matrix and the weight matrix is the right matrix. Exemplarily, still taking the 4 GPUs in the previous example as an example, the operation result obtained by each GPU in the current operation stage is a data matrix C3 with a size of (m, k). Then, each GPU can perform the division operation on the data matrix C3 with column as the splitting dimension to obtain 4 second data sub-matrices with a size of (m, k / 4):
[0126] The second data sub-matrix C 30 : corresponding to the 1st to k / 4th columns of C3;
[0127] The second data sub-matrix C 31 : corresponding to the (k / 4 + 1)th to k / 2th columns of C3;
[0128] The second data sub - matrix C 32 : corresponding to the columns from \(k / 2 + 1\) to \(3k / 4\) of C3;
[0129] The second data sub - matrix C 33 : corresponding to the columns from \(3k / 4+1\) to \(k\) of C3.
[0130] Meanwhile, for the weight matrix B3 (with size \((k,n)\)) in the current operation stage, a partitioning operation is also performed with the row as the splitting dimension, obtaining 4 weight sub - matrices with size \((k / 4,n)\):
[0131] The weight sub - matrix B stored in GPU0 30 : corresponding to the rows from 1 to \(k / 4\) of B3;
[0132] The weight sub - matrix B stored in GPU1 31 : corresponding to the rows from \(k / 4 + 1\) to \(k / 2\) of B3;
[0133] The weight sub - matrix B stored in GPU2 32 : corresponding to the rows from \(k / 2 + 1\) to \(3k / 4\) of B3;
[0134] The weight sub - matrix B stored in GPU3 33 : corresponding to the rows from \(3k / 4 + 1\) to \(k\) of B3.
[0135] Alternatively, in another example, during the multiplication operation in the next execution stage, the data matrix is the right matrix and the weight matrix is the left matrix. Exemplarily, still taking the 4 GPUs in the previous example, the operation result obtained by each GPU in the current running stage is a data matrix C3 with size \((k,n)\). Then each GPU can perform a partitioning operation on the data matrix C3 with the row as the splitting dimension, obtaining 4 second data sub - matrices with size \((k / 4,n)\):
[0136] The second data sub - matrix C 30 : corresponding to the rows from 1 to \(k / 4\) of C3;
[0137] The second data sub - matrix C 31 : corresponding to the rows from \(k / 4 + 1\) to \(k / 2\) of C3;
[0138] The second data sub - matrix C 32 : corresponding to the rows from \(k / 2 + 1\) to \(3k / 4\) of C3;
[0139] The second data sub - matrix C 33 : corresponding to the rows from \(3k / 4 + 1\) to \(k\) of C3.
[0140] Meanwhile, for the weight matrix B3 (with size (m, k)) in the current operation stage, a partitioning operation is also performed with columns as the partitioning dimension, resulting in 4 weight sub-matrices with size (m, k / 4):
[0141] The weight sub-matrix B stored in GPU0 30 : corresponding to the 1st to k / 4th columns of B3;
[0142] The weight sub-matrix B stored in GPU1 31 : corresponding to the (k / 4 + 1)th to k / 2th columns of B3;
[0143] The weight sub-matrix B stored in GPU2 32 : corresponding to the (k / 2 + 1)th to 3k / 4th columns of B3;
[0144] The weight sub-matrix B stored in GPU3 33 : corresponding to the (3k / 4 + 1)th to kth columns of B3.
[0145] Furthermore, C 30 can be retained in the memory resources of GPU0 itself. In the next execution stage, the operation with B 30 is performed within GPU0, and C 31 is retained in the memory resources of GPU1 itself. In the next execution stage, the operation with B 31 is performed within GPU1, C 32 is retained in the memory resources of GPU2 itself. In the next execution stage, the operation with B 32 is performed within GPU2, C 33 is retained in the memory resources of GPU3 itself. In the next execution stage, the operation with B 33 is performed within GPU3, and then the operation results C 40 、C 41 、C 42 and C 43 with size (m, n) are obtained respectively. Among them, C 40 、C 41 、C 42 and C 43 are aggregated and combined, which is the final operation result C4 of the operator in the next execution stage.
[0146] In this way, in the case of the conventional parallel strategy corresponding to the current execution stage, based on the parallel strategy corresponding to the next execution stage, the data required for the operator operation in the next execution stage can be directly written into the memory resources of each GPU, thereby reducing the data volume transmitted between GPUs, reducing the communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0147] In an exemplary embodiment, the parallel strategy corresponding to the execution stage is a data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, a part or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip, including at least one of the following:
[0148] If the parallel strategy corresponding to the next execution stage is a conventional parallel strategy or a tensor parallel strategy, for any artificial intelligence chip, the operation results of the artificial intelligence chip in the execution stage are written into the memory resources of each artificial intelligence chip respectively;
[0149] If the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy, for any artificial intelligence chip, according to the number of artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, taking the column as the division dimension, the operation results of the artificial intelligence chip in the execution stage are divided into multiple third data sub-matrices corresponding to each artificial intelligence chip respectively, and the multiple third data sub-matrices are written into the memory resources corresponding to their respective artificial intelligence chips respectively; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, the operation results of the artificial intelligence chip in the execution stage are retained in the memory resources of the artificial intelligence chip itself.
[0150] In the embodiment of the present application, it is assumed that the current execution stage corresponds to the third execution stage, and its corresponding parallel strategy is a data parallel strategy. The data parallel strategy means that the data matrix A is divided into multiple parts and written into the memory resources of the corresponding GPUs respectively. Inside each GPU, it performs operations with the full weight matrix B1. After the operation is completed, the operation results of each GPU are part of the operation results of the operator in the current execution stage. By summarizing and combining the operation results of each GPU, the operation results of the operator in the current execution stage can be obtained.
[0151] Exemplarily, in the previous example, GPU0, GPU1, GPU2, and GPU3 respectively obtain the operation results C 20 、C 21 、C 22 and C 23 as an example.
[0152] In one example, the parallel strategy corresponding to the next execution stage is a tensor parallel strategy or a conventional parallel strategy. Since both the tensor parallel strategy and the conventional parallel strategy use the complete data matrix to perform operations with part or all of the weight matrices stored inside the GPU, and the operation results of each GPU in the current execution stage (data parallel stage) are part of the complete operation results, data synchronization is required between each GPU at this time to write their respective operation results into their own memory resources and the memory resources of other GPUs.
[0153] For example, keep C 20 in the memory resources of GPU0 and write it into the memory resources of GPU1, GPU2, and GPU3 respectively. Keep C 21 in the memory resources of GPU1 and write it into the memory resources of GPU0, GPU2, and GPU3 respectively. Keep C 22 in the memory resources of GPU2 and write it into the memory resources of GPU0, GPU1, and GPU3 respectively. Keep C 23 in the memory resources of GPU3 and write it into the memory resources of GPU0, GPU1, and GPU2 respectively. In this way, each GPU can obtain C 20 , C 21 , C 22 , and C 23 . After summarizing and merging, the complete operation result C2 can be obtained. Then, after entering the next execution stage, it can perform operations with some or all of the weight matrices stored inside each GPU to complete tensor parallel or regular parallel operations and obtain the corresponding operation results.
[0154] In another example, the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy. The hybrid parallel strategy divides the complete data matrix into multiple data matrices by columns or rows according to the number of GPUs, and divides the complete weight matrix into multiple weight matrices by rows or columns and stores them in each GPU. For example, divide the weight matrix W with a shape of (k, p) into 4 weight sub-matrices W0, W1, W2, and W3 evenly by rows, and the shape of each sub-matrix is (k / 4, p), and store them in each GPU respectively. Then write the multiple data matrices divided by columns into the memory resources of each GPU to perform operations with the partial weight matrices stored in them; or divide the weight matrix W with a shape of (p, k) into 4 weight sub-matrices W0, W1, W2, and W3 evenly by columns, and the shape of each sub-matrix is (p, k / 4), and store them in each GPU respectively. Then write the multiple data matrices divided by rows into the memory resources of each GPU to perform operations with the partial weight matrices stored in them.
[0155] Since after the operations in the current execution stage (data parallel stage) are completed, the operation results of each GPU respectively correspond to partial rows of the complete operation result, and the final operation result of the operator in the current execution stage is obtained after merging. When entering the next execution stage, since the next execution stage is a hybrid parallel strategy, if the data matrix in the next execution stage is the left matrix and the weight matrix is the right matrix, then each GPU will divide and obtain partial columns of the final operation result. Taking the scale of the operation result of each GPU as (m / 4, k) and the final operation result as (m, k) as an example, each GPU in the next execution stage will divide and obtain a data matrix with a scale of (m, k / 4).
[0156] Therefore, after the operation in the current execution stage is completed, after dividing the operation results of each GPU by columns to obtain multiple third data sub-matrices, each third data sub-matrix can be written into the memory resources of each GPU respectively, and spliced by rows. Each GPU can obtain partial columns of the final operation result in the current execution stage, and then enter the next execution stage. In each GPU, it can directly operate with the partial weight matrix obtained by row division stored to obtain the corresponding operation result.
[0157] Taking the 4 GPUs in the previous example as an example, the operation result C of each GPU 20 、C 21 、C 22 and C 23 all perform the operation of evenly dividing by columns into 4 third data sub-matrices to obtain the third data sub-matrices G i0 、G i1 、G i2 、G i3 , where i is the index of the GPU (the value range is 0, 1, 2, 3), and the scale of each third data sub-matrix is (m / 4, k / 4), where:
[0158] The third data sub-matrix G i0 corresponds to the 1st to k / 4th columns, the third data sub-matrix G i1 corresponds to the (k / 4 + 1)th to k / 2th columns, the third data sub-matrix G i2 corresponds to the (k / 2 + 1)th to 3k / 4th columns, and the third data sub-matrix G i3 corresponds to the (3k / 4 + 1)th to kth columns.
[0159] Write G ij directly into the memory resources of GPU j , then GPU j will obtain 4 third data sub-matrices with a scale of (m / 4, k / 4). After splicing by rows, a data sub-matrix with a scale of (m, k / 4) can be obtained, where j is the index of the GPU (the value range is 0, 1, 2, 3). When entering the operation of the next execution stage, the spliced data sub-matrix is operated with the weight sub-matrix obtained by row division inside the GPU, and an operation result with a scale of (m, p) can be obtained. By performing allreduce processing on the operation results of each GPU, the final operation result corresponding to the operator in this running stage can be obtained.
[0160] Alternatively, if the data matrix in the next execution stage is the right matrix and the weight matrix is the left matrix, each GPU will divide and obtain partial rows of the final operation result. Taking the operation result scale of each GPU as (k / 4, m) and the final operation result as (k, m) as an example, the data matrix with a scale of (k / 4, m) will be divided and obtained by each GPU in the next execution stage. Therefore, after the operation in the current execution stage is completed, the operation results of each GPU can be retained in the memory resources of each GPU itself.
[0161] Exemplarily, taking GPU0, GPU1, GPU2, and GPU3 as examples, which respectively obtain operation results C with a scale of (k / 4, m) 20 、C 21 、C 22 and C 23 as an example, then C 20 can be retained in the memory resources of GPU0 and operated with the weight sub-matrix W0 with a scale of (p, k / 4) stored in GPU0; C 21 can be retained in the memory resources of GPU1 and operated with the weight sub-matrix W1 with a scale of (p, k / 4) stored in GPU1; C 22 can be retained in the memory resources of GPU2 and operated with the weight sub-matrix W2 with a scale of (p, k / 4) stored in GPU2; C 23 can be retained in the memory resources of GPU3 and operated with the weight sub-matrix W3 with a scale of (p, k / 4) stored in GPU3. In this way, each GPU can obtain partial rows of the final operation result in the current execution stage, and thus enter the next execution stage, where it can directly operate with the partial weight matrix obtained by column division and stored in each GPU to obtain the corresponding operation result.
[0162] The method provided in the embodiments of the present application, in the case of corresponding data parallel strategies in the current execution stage, can directly write the data required for the operator operation in the next execution stage into the memory resources of each GPU based on the corresponding parallel strategy in the next execution stage, thereby reducing the amount of data transmitted between GPUs, reducing communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0163] In an exemplary embodiment, the corresponding parallel strategy in the execution stage is the tensor parallel strategy. Based on the corresponding parallel strategy in the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0164] If the corresponding parallel strategy in the next execution stage is the conventional parallel strategy, for any artificial intelligence chip, write the operation results of the artificial intelligence chip in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0165] If the parallel strategy corresponding to the next execution stage is a data parallel strategy, then for any artificial intelligence chip, according to the number of artificial intelligence chips, with the row division dimension, the operation results of the artificial intelligence chip in the execution stage are divided into multiple fourth data sub-matrices respectively corresponding to each artificial intelligence chip, and the multiple fourth data sub-matrices are respectively written into the memory resources of their corresponding artificial intelligence chips;
[0166] If the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy, then for any artificial intelligence chip, when the data matrix in the next execution stage is the left matrix of the multiplication operation, the operation results of the artificial intelligence chip in the execution stage are retained in the memory resources of the artificial intelligence chip itself; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with the row division dimension, the operation results of the artificial intelligence chip in the execution stage are divided into multiple fifth data sub-matrices respectively corresponding to each artificial intelligence chip, and the multiple fifth data sub-matrices are respectively written into the memory resources of their corresponding artificial intelligence chips.
[0167] In the embodiments of the present application, it is assumed that the current execution stage corresponds to the second execution stage, and the corresponding parallel strategy is a tensor parallel strategy. The tensor parallel strategy means that the weight matrix B is divided into multiple parts and stored in the corresponding GPUs respectively, and the operation is performed with the full data matrix in each GPU. After the operation is completed, the operation results of each GPU are part of the operation results of the operator in the current execution stage. By summarizing and combining the operation results of each GPU, the operation results of the operator in the current execution stage can be obtained.
[0168] Exemplarily, still taking the 4 GPUs, GPU0, GPU1, GPU2, and GPU3 in the foregoing example as an example, the operation results obtained by GPU0, GPU1, GPU2, and GPU3 in the current execution stage (tensor parallel stage) are C with a size of (m, q / 4) 50 、C 51 、C 52 and C 53 respectively. Taking C 50 、C 51 、C 52 and C 53 as an example, the summary and combination of C
[0169] In one example, the parallel strategy corresponding to the next execution stage is a conventional parallel strategy. Since the conventional parallel strategy performs operations using the complete data matrix and the complete weight matrix, and the operation results of each GPU in the current execution stage are part of the complete operation result C5, data synchronization is required among the GPUs at this time to write their respective operation results into their own memory resources and the memory resources of other GPUs. In this way, each GPU can obtain the operation result C 50 、C 51 、C 52 and C 53 . After summarizing and merging, the complete operation result C5 can be obtained, and then operations can be performed within each GPU with the complete weight matrix stored therein to obtain the corresponding operation results.
[0170] In another example, the parallel strategy corresponding to the next execution stage is a data parallel strategy. The data parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs, and then writes the multiple data matrices into the memory resources of each GPU respectively for operations with the full weight matrix. That is, in the next execution stage, each GPU operates on a part of the rows in the complete operation result of the operator in the current execution stage.
[0171] Since the operation result obtained by each GPU in the current execution stage (tensor parallel stage) is a part of the columns in the complete operation result of the operator in the current execution stage, the operation results of each GPU can be directly divided into multiple fourth data sub-matrices corresponding to each GPU respectively with the row as the division dimension.
[0172] Taking the 4 GPUs in the previous example as an example, the operation result C 50 、C 51 、C 52 and C 53 of each GPU all perform the operation of evenly dividing into 4 fourth data sub-matrices by rows to obtain the fourth data sub-matrices R i0 、R i1 、R i2 、R i3 , where i is the index of the GPU (taking values 0, 1, 2, 3), and the scale of each fourth data sub-matrix is (m / 4, k / 4), where:
[0173] The fourth data sub-matrix R i0 corresponds to columns 1 to m / 4, the fourth data sub-matrix R i1 corresponds to columns m / 4 + 1 to m / 2, the fourth data sub-matrix R i2 corresponds to columns m / 2 + 1 to 3m / 4, and the fourth data sub-matrix R i3 corresponds to columns 3m / 4 + 1 to m.
[0174] Write R ij directly into the memory resources of the GPU j , then the GPU j will obtain four fourth data sub - matrices with dimensions (m / 4, k / 4). After concatenating them by columns, a data sub - matrix with dimensions (m / 4, k) can be obtained. Here, j is the index of the GPU (taking values 0, 1, 2, 3). When entering the operations of the next execution stage, the concatenated data sub - matrix is operated on with the complete weight matrix (k, p) stored inside the GPU, and an operation result with dimensions (m / 4, p) can be obtained. By aggregating and combining the operation results of each GPU, the final operation result with dimensions (m, p) corresponding to the operator in this execution stage can be obtained.
[0175] In another example, the parallel strategy corresponding to the next execution stage is a hybrid parallel strategy. The hybrid parallel strategy is to divide the complete weight matrix into multiple weight matrices and store them in each GPU. For example: when the data matrix is the left matrix and the weight matrix is the right matrix, the weight matrix W with shape (k, p) can be evenly divided into 4 weight sub - matrices W0, W1, W2, W3 by rows, each sub - matrix with shape (k / 4, p), and stored in each GPU respectively. Then, according to the number of GPUs, the complete data matrix is divided into multiple data matrices by columns, and these multiple data matrices are written into the memory resources of each GPU respectively to be operated on with the partial weight matrix stored in it; or, when the data matrix is the right matrix and the weight matrix is the left matrix, the weight matrix W with shape (p, k) can be evenly divided into 4 weight sub - matrices W0, W1, W2, W3 by columns, each sub - matrix with shape (p, k / 4), and stored in each GPU respectively. Then, according to the number of GPUs, the complete data matrix is divided into multiple data matrices by rows, and these multiple data matrices are written into the memory resources of each GPU respectively to be operated on with the partial weight matrix stored in it.
[0176] In one example, the data matrix in the next execution stage is the left matrix for multiplication. Since after the current execution stage (tensor parallel stage), the operation result obtained by each GPU is a partial column in the complete operation result of the operator in the current execution stage, with dimensions (m, k / 4), it can no longer be further divided at this time and is directly retained in the memory resources of the GPU itself. In the next execution stage, this operation result is used to operate on the weight sub - matrix stored inside the GPU, and an operation result with dimensions (m, p) can be obtained. By performing an all - reduce operation on the current operation results of each GPU with dimensions (m, p), the final operation result corresponding to the operator in this execution stage can be obtained.
[0177] In another example, the data matrix in the next execution stage is the right matrix for the multiplication operation. Since after the completion of the current execution stage (tensor parallelism stage), the operation result obtained by each GPU is a partial column in the complete operation result of the operator in the current execution stage, with a size of (m, k / 4), at this time, the operation result can be further divided by rows, and the fifth data sub-matrices with a size of (m / 4, k / 4) obtained by the division are respectively written into the memory resources of the corresponding GPUs. After each GPU merges the multiple written fifth data sub-matrices by columns, the GPU can obtain a data matrix with a size of (m / 4, k) (i.e., partial rows of the final operation result in the current execution stage). In the next execution stage, this data matrix is used to perform an operation with the weight sub-matrix stored inside the GPU, and full reduction processing is performed on the operation results of each GPU, then the final operation result corresponding to the operator in this operation stage can be obtained. For the specific process, refer to the relevant description in the foregoing embodiments, and it will not be elaborated herein in the embodiments of the present application.
[0178] In this way, in the case of the tensor parallelism strategy corresponding to the current execution stage, based on the parallelism strategy corresponding to the next execution stage, the data required for the operator operation in the next execution stage can be directly written into the memory resources of each GPU, thereby reducing the amount of data transferred between GPUs, reducing the communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0179] In an exemplary embodiment, the parallelism strategy corresponding to the execution stage is a hybrid parallelism strategy. Based on the parallelism strategy corresponding to the next execution stage, some or all of the operation results of each artificial intelligence chip in the execution stage are directly written into the memory resources corresponding to the target artificial intelligence chip, including:
[0180] Perform full reduction processing on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction processing result:
[0181] If the parallelism strategy corresponding to the next execution stage is a tensor parallelism strategy or a conventional parallelism strategy, then for any artificial intelligence chip, retain the full reduction processing result in the memory resources of the artificial intelligence chip itself;
[0182] If the parallelism strategy corresponding to the next execution stage is a data parallelism strategy, then for any artificial intelligence chip, according to the number of artificial intelligence chips, with the row as the division dimension, divide the full reduction processing result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself.
[0183] In the embodiment of the present application, it is assumed that the current execution stage corresponds to the fourth execution stage, and its corresponding parallel strategy is a hybrid parallel strategy. The hybrid parallel strategy can be referred to the description of the foregoing embodiment, and will not be elaborated herein in the embodiment of the present application. After each GPU performs operations using the hybrid parallel strategy, it is necessary to perform a global reduction process on the operation results of each artificial intelligence chip in the execution stage to obtain the global reduction process result.
[0184] In one example, the parallel strategy corresponding to the next execution stage is a tensor parallel strategy or a conventional parallel strategy. Since both the tensor parallel strategy and the conventional parallel strategy operate on a complete data matrix, the global reduction process result is a complete data matrix. Therefore, for any GPU, there is no need to perform data synchronization, and the global reduction process result can be directly retained in the memory resources of the GPU itself and operated with the complete or partial weight matrix stored in the GPU.
[0185] In another example, the parallel strategy corresponding to the next execution stage is a data parallel strategy. Since the data parallel strategy divides the complete data matrix into multiple data matrices according to the number of GPUs, and then writes the multiple data matrices into the memory resources of each GPU respectively to operate with the full amount of weight matrix. Since the global reduction process result is a complete data matrix, the global reduction process result can be directly divided into multiple sixth data sub-matrices with the row as the division dimension. Each sixth data sub-matrix corresponds to a GPU, and the sixth data sub-matrix corresponding to each GPU itself can be retained in the memory resources of the GPU itself. The specific process can refer to the division and direct writing operation of the foregoing first data sub-matrix.
[0186] In this way, in the case where the current execution stage corresponds to the hybrid parallel strategy, based on the parallel strategy corresponding to the next execution stage, the data required for the operation of the next execution stage operator can be directly written into the memory resources of each GPU, thereby reducing the amount of data transmitted between GPUs, reducing the communication time consumption, and then significantly improving the inference efficiency of the large model and reducing the inference latency.
[0187] To enable those skilled in the art to better understand the embodiments of the present application, the embodiments of the present application are described below through specific examples.
[0188] Exemplarily, taking Figure 1 the attention operator flow shown as an example, using the operation method of the operator flow provided by the embodiment of the present application, the operators in the attention operator flow are divided into four execution stages according to the matrix calculation scale and characteristics, as shown in reference to Figure 3 shown.
[0189] Among them, the data matrices and weight matrices of the operations of the first MMA operator, the split-root mean square normalization operator, and the rotary position embedding operator executed by tensor kv_c are all small, so they are divided into the same execution stage, that is, the first execution stage, and this execution stage corresponds to the conventional parallel strategy. Figure 3 In it, the operations of the first MMA operator, the split-root mean square normalization operator, and the rotary position embedding operator executed by tensor kv_c are fused into a matrix multiply-accumulate-root mean square normalization-update cache operator (which can also be expressed as MMA_rmsnorm_updateCache).
[0190] The data matrix of the operation of the second MMA operator and the position embedding operator executed by tensor q_c is small, but the weight matrix is large, so they are divided into the same execution stage, that is, the second execution stage, and this execution stage corresponds to the tensor parallel strategy. Figure 3 In it, the second MMA operator and the position embedding operator executed by tensor q_c are fused into a matrix multiply-accumulate-query embedding operator (which can also be expressed as MMA_query_embedding).
[0191] The data matrices of the operations of the third MMA operator, the softmax operator, and the fourth MMA operator are large, but the weight matrices are small, so they are divided into the same execution stage, that is, the third execution stage, and this execution stage corresponds to the data parallel strategy. Figure 3 In it, the third MMA operator, the softmax operator, and the fourth MMA operator are fused into a grouped query attention operator (Grouped Query Attention, GQA).
[0192] The data matrix and weight matrix of the attention multiplication matrix operator are both large, and it is used as an execution stage, that is, the fourth execution stage, and this execution stage corresponds to the hybrid parallel strategy.
[0193] The input data enters the first execution stage. Since the first execution stage corresponds to the conventional parallel strategy, that is, the input data is written into each GPU, and the operation of the input data and the complete weight matrix is performed in each GPU, obtaining an intermediate tensor of shape (B, 1536), and an intermediate tensor of shape (B, S, 576) (it should be noted that this intermediate tensor is not shown in Figure 3 ).
[0194] Since the second execution stage adopts the tensor parallel strategy, for the operation results of each GPU in the first execution stage, the operation results (tensor q_c and tensor kv) can be directly retained in the memory resources of each GPU, enter the second execution stage, and perform tensor parallel operations on the written data matrix and part of the weight matrix inside each GPU.
[0195] After completing the tensor parallel operation in the second execution stage, each GPU obtains an intermediate tensor with the shape of (B, 128×576 / 4). After aggregating and combining the operation results of each GPU, the operation result of the operator in the second execution stage is obtained. Since the data parallel strategy is adopted in the third execution stage, for the operation results of each GPU in the second execution stage, the operation results can be split into 4 data matrices corresponding to 4 GPUs respectively according to rows (with the shape of (B / 4, 128×576 / 4)), and the split data matrices are respectively written into the memory resources of their corresponding GPUs ( Figure 3 Only the splitting schematic of the operation result of one GPU is shown in
[0196] for reference in splitting the operation results of other GPUs), then at this time each GPU obtains partial rows of the operation result of the operator in the second execution stage, enters the third execution stage, and performs data parallel operation on the partial data matrix written in each GPU and the complete weight matrix inside each GPU. Figure 3 After completing the data parallel operation in the third execution stage, each GPU obtains an intermediate tensor with the shape of (B / 4, 128×576). Since the hybrid parallel strategy is adopted in the fourth stage, for the operation results of each GPU in the third execution stage, the operation results can be split into 4 data matrices corresponding to 4 GPUs respectively according to columns (with the shape of (B / 4, 128×512 / 4)), and the split data matrices are respectively written into the memory resources of their corresponding GPUs (
[0197] Only the splitting schematic of the operation result of one GPU is shown in
[0198] for reference in splitting the operation results of other GPUs), then at this time each GPU obtains partial columns of the operation result of the operator in the third execution stage, enters the fourth execution stage, and performs data parallel operation on the partial data matrix written in each GPU and the partial weight matrix stored inside each GPU.
[0199] After completing the hybrid parallel operation in the fourth execution stage, perform allreduce on the operation results output by the four GPUs to obtain the final result. The parallel operation method of the operator stream provided by the embodiment of the present application can flexibly select a parallel strategy according to the calculation characteristics of the operator stream, and after the calculation is completed, according to the requirements of the subsequent stage, write part or all of the operation results directly to the memory resources of the corresponding GPU card, giving full play to the operator performance, and can not only meet the better latency performance requirements, but also have better throughput.In addition, it should be noted that the parallel operation method of the operator flow provided in the embodiments of the present application can be applied to the inference process in at least fields such as speech processing, image processing, text processing, and video processing. Among them, in the field of speech processing, the data matrices of the operations of the operators in the target operator flow can be audio data or audio feature data; in the field of image processing, the data matrices of the operations of the operators in the target operator flow can be image data or image feature data; in the field of text processing, the data matrices of the operations of the operators in the target operator flow can be text data or text feature data; in the field of video processing, the data matrices of the operations of the operators in the target operator flow can be video data or video feature data, etc.
[0200] In the field of image recognition, for complex large model inference tasks, a large amount of computation and data transmission are often involved, especially when processing models based on the attention mechanism. The application of this solution in the image recognition scenario will be elaborated in detail below by using 4 GPUs to accelerate the inference process of the attention operator flow.
[0201] Suppose a batch of image data needs to be recognized to detect target objects therein, for example: whether there are specific geographical features such as airports, ports, etc.; or whether there are people or items with specific features. In this example, a large model based on the Transformer architecture is adopted, where the attention mechanism is the core component of the model, and the target operator flow includes the attention mechanism and multiple related operators before and after it.
[0202] First, according to the matrix scale of the operations of the operators in the target operator flow and the dependencies between them, the target operator flow is divided into multiple execution stages, and a suitable parallel strategy is selected for each execution stage.
[0203] For example: The operators for feature extraction and preliminary transformation are divided into the first execution stage. The operators in this stage are involved in feature extraction of the input image and preliminary linear transformation, with relatively small data matrices and large weight matrices. For example, a convolution operation is performed on the input image block to map the image features to a higher-dimensional space. For the first execution stage, the tensor parallel strategy is adopted, and the weight matrix is divided among multiple GPUs, and each GPU is responsible for calculating the product of a part of the weight matrix and the input data.
[0204] The operators used for attention score calculation are divided into the second execution stage. The operators in this stage are mainly used to calculate attention scores. The data matrix is relatively large because it contains the image features after preliminary transformation, while the weight matrix is relatively small. For the second execution stage, a data parallel strategy can be adopted, evenly distributing the input image feature data to 4 GPUs, and each GPU independently calculates the attention scores for the data part it is responsible for.
[0205] During the execution of the target operator stream, in the first execution stage, each GPU performs matrix multiplication operations on the allocated part of the weight matrix and the input image data according to the tensor parallel strategy. For example, GPU1 is responsible for calculating the product of the first part of the weight matrix and the input image data, GPU2 is responsible for the second part, and so on.
[0206] When the first execution stage is completed, according to the data parallel strategy adopted in the second execution stage, the intermediate results calculated by each GPU (i.e., the image features after preliminary transformation) are divided into multiple parts by rows according to the corresponding partitioning method of the data parallel strategy, and each part is written into the memory resources of its corresponding GPU (including itself and other GPUs). Finally, each GPU has the complete image features for calculating attention scores.
[0207] Each GPU calculates the attention scores for the allocated image feature data according to the data parallel strategy. Each GPU independently completes the calculation task without excessive communication. When the second execution stage is completed, the attention score results can also be directly written into the memory resources of the corresponding GPU according to the parallel strategy of the subsequent execution stage, preparing for subsequent calculations.
[0208] In the above way, the calculation tasks of each execution stage in the target operator stream are completed in sequence. Each stage selects an appropriate parallel strategy according to its matrix calculation scale and dependency relationship, and performs efficient data transmission between stages. After all execution stages are completed, the final operation results of the 4 GPUs are integrated to obtain the final image recognition result. For example, by summarizing and analyzing the attention weights calculated by each GPU, it is judged whether there are specific geographical features in the image.
[0209] By dividing the target operator stream into multiple execution stages, selecting appropriate parallel strategies according to the matrix calculation scale and dependency relationship, and performing efficient data transmission between different stages, the parallel requirements of each operator can be accurately adapted, the communication time consumption can be reduced, and thus the inference efficiency of the large model can be significantly improved. In the field of image recognition, this method can process a large amount of high-definition image data faster, meeting the requirements of real-time and accuracy.
[0210] Embodiments of the present application can be implemented through a computing cluster including multiple artificial intelligence chips. Refer to Figure 4 As shown, a computing cluster including multiple GPUs is shown. Among them, the GPU contains video memory for storing data. There are multiple SPCs (Streaming Processing Clusters) in the GPU, and each SPC contains computing units. There is on-chip cache in the computing units for quickly storing and reading data to accelerate the computing process. The computing units also contain multiple execution units, and components such as physical registers are provided in the execution units, which can be used to temporarily store data during the computing process, etc.
[0211] The operator operations in each execution stage can be undertaken by the computing units of the GPU. For example, in the execution stage of calculating the attention score, the execution units in the computing units of each GPU are responsible for specific matrix multiplication and other operation operations under the data parallel strategy. The physical registers are used to temporarily store the intermediate calculation data, and the on-chip cache speeds up the data reading and storage, improving the operation speed.
[0212] The video memory can be used to store the input data, intermediate results, and final results of each execution stage. For example, after the operation in Execution Stage 1 is completed, the intermediate result can be temporarily stored in the video memory, and then according to the data parallel strategy of the next execution stage, it is read from the video memory and directly written to the corresponding GPU. The SPC plays a role in coordinating the data transfer between the video memory and the computing units, ensuring that the data can be accurately and timely supplied to the computing units for subsequent operator operations.
[0213] Through the collaborative work of the components in the GPU and in cooperation with the parallel strategies of different execution stages, the operation of the operator stream can be completed more efficiently, reducing the data waiting time and unnecessary communication overhead, and providing support for improving the inference efficiency of large models from the hardware level.
[0214] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0215] Based on the same inventive concept, an embodiment of this application further provides a parallel operation device for an operator flow for implementing the parallel operation method of the operator flow involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the parallel operation device for an operator flow provided below can refer to the limitations on the parallel operation method of the operator flow in the above text, and will not be elaborated here.
[0216] In an exemplary embodiment, as Figure 5 shown, a parallel operation device 500 for an operator flow is provided, including: a partitioning module 502, an execution module 504, and a determination module 506, where:
[0217] The partitioning module 502 is configured to partition a target operator flow into multiple execution stages, each execution stage includes at least one operator, and the operators belonging to the same execution stage correspond to the same parallel strategy, and the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy;
[0218] The execution module 504 is configured to, during the running process of the target operator flow, for any one of the execution stages, adopt the parallel strategy corresponding to the execution stage inside each artificial intelligence chip, execute the operations corresponding to the operators in the execution stage, and after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, directly write part or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip, and the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips;
[0219] The determination module 506 is configured to, after all execution stages are completed, obtain the operation result of the target operator flow based on the final operation results of each artificial intelligence chip.
[0220] By using the parallel operation device for an operator flow provided by the embodiment of this application, during the execution process of the target operator flow, the corresponding parallel strategy can be accurately adapted according to the operator characteristics of the operators in the target operator flow, and at the same time, based on the parallel strategies adopted in different execution stages, the direct writing method of the operation results between different artificial intelligence chips is finely optimized, so that the operators can be fully utilized, the calculation speed of the target operator flow can be significantly improved, and further the inference latency of the large model can be reduced and the inference efficiency of the large model can be improved.
[0221] In one of the embodiments, the conventional parallel strategy includes that each artificial intelligence chip executes the operation corresponding to the operator on the full amount of data matrix and the full amount of weight matrix;
[0222] The hybrid parallel strategy includes: when the data matrix is the left matrix in a multiplication operation, splitting the data matrix by columns into multiple data sub-matrices, and splitting the weight matrix by rows into multiple weight sub-matrices; or, when the data matrix is the right matrix in a multiplication operation, splitting the data matrix by rows into multiple data sub-matrices, and splitting the weights by columns into multiple weight sub-matrices; each artificial intelligence chip performs the operation corresponding to the operator based on the allocated data sub-matrix and weight sub-matrix.
[0223] In one embodiment, the execution phase includes at least one of the following:
[0224] The first execution phase, the first execution phase includes at least one first operator, the first operator is an operator for which neither the data matrix nor the weight matrix of the operation satisfies the parallel requirement, and the first execution phase corresponds to the conventional parallel strategy;
[0225] The second execution phase, the second execution phase includes at least one second operator, the second operator is an operator for which the data matrix of the operation does not satisfy the parallel requirement, but the weight matrix of the operation satisfies the parallel requirement, and the second execution phase corresponds to the tensor parallel strategy;
[0226] The third execution phase, the third execution phase includes at least one third operator, the third operator is an operator for which the data matrix of the operation satisfies the parallel requirement, but the weight matrix of the operation does not satisfy the parallel requirement, and the third execution phase corresponds to the data parallel strategy;
[0227] The fourth execution phase, the fourth execution phase includes at least one fourth operator, the fourth operator is an operator for which both the data matrix and the weight matrix of the operation satisfy the parallel requirement, and the fourth execution phase corresponds to the hybrid parallel strategy.
[0228] In one embodiment, the parallel strategy corresponding to the execution phase is the conventional parallel strategy. Based on the parallel strategy corresponding to the next execution phase, writing some or all of the operation results of each artificial intelligence chip in the execution phase directly into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0229] If the parallel strategy corresponding to the next execution phase is the tensor parallel strategy, then for any artificial intelligence chip, retain the operation result of the artificial intelligence chip in the execution phase in the memory resource of the artificial intelligence chip itself;
[0230] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with the row division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple first data sub-matrices respectively corresponding to each artificial intelligence chip, and retain the first data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself;
[0231] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices respectively corresponding to each artificial intelligence chip, and retain the second data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself, where the division method of the second data sub-matrix includes:
[0232] In the case where the data matrix in the next execution stage is the left matrix of the multiplication operation, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices respectively corresponding to each artificial intelligence chip with the column division dimension; or, in the case where the data matrix in the next execution stage is the right matrix of the multiplication operation, divide the operation results of the artificial intelligence chips in the execution stage into multiple second data sub-matrices respectively corresponding to each artificial intelligence chip with the row division dimension.
[0233] In one embodiment, the parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resource corresponding to the target artificial intelligence chip includes at least one of the following:
[0234] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0235] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, according to the number of artificial intelligence chips, with columns as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple third data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves.
[0236] In one embodiment, the parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following:
[0237] If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively;
[0238] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fourth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips;
[0239] If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fifth data sub-matrices respectively corresponding to each artificial intelligence chip, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
[0240] In one embodiment, the parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing some or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes:
[0241] Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and execute at least one of the following for the full reduction process result:
[0242] If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, for any artificial intelligence chip, retain the full reduction process result in the memory resources of the artificial intelligence chip itself;
[0243] If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any artificial intelligence chip, divide the full reduction process result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip in the row division dimension according to the number of artificial intelligence chips, and retain the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself.
[0244] Each module in the parallel operation device of the above operator flow can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0245] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it realizes a method for parallel operation of an operator flow. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0246] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0247] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are realized.
[0248] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are realized.
[0249] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are realized.
[0250] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0251] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0252] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0253] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A parallel running method for an operator stream, characterized in that The method includes: Dividing a target operator stream into multiple execution stages, each execution stage including at least one operator, and the operators belonging to the same execution stage corresponding to the same parallel strategy, where the parallel strategy includes at least one of a conventional parallel strategy, a data parallel strategy, a tensor parallel strategy, and a hybrid parallel strategy; During the running of the target operator stream, for any one of the execution stages, adopting the parallel strategy corresponding to the execution stage inside each artificial intelligence chip to execute the operations corresponding to the operators in the execution stage, and after the operations in the execution stage are completed, based on the parallel strategy corresponding to the next execution stage, directly writing part or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip, where the target artificial intelligence chip includes each artificial intelligence chip itself or all artificial intelligence chips; After the calculations in each execution stage are completed, obtaining the operation result of the target operator stream based on the final operation results of each artificial intelligence chip; Among them, the conventional parallel strategy includes that each artificial intelligence chip executes the operation corresponding to the operator on the full data matrix and the full weight matrix; The hybrid parallel strategy includes: when the data matrix is the left matrix of a multiplication operation, splitting the data matrix into multiple data sub-matrices with columns as the splitting dimension and splitting the weight matrix into multiple weight sub-matrices with rows as the splitting dimension; or when the data matrix is the right matrix of a multiplication operation, splitting the data matrix into multiple data sub-matrices with rows as the splitting dimension and splitting the weight matrix into multiple weight sub-matrices with columns as the splitting dimension; each artificial intelligence chip executes the operation corresponding to the operator according to the allocated data sub-matrix and weight sub-matrix; The execution stage includes at least one of the following: The first execution stage, the first execution stage including at least one first operator, where the first operator is an operator for which neither the data matrix nor the weight matrix of the operation meets the parallel requirement, and the first execution stage corresponds to the conventional parallel strategy; The second execution stage, the second execution stage including at least one second operator, where the second operator is an operator for which the data matrix of the operation does not meet the parallel requirement but the weight matrix of the operation meets the parallel requirement, and the second execution stage corresponds to the tensor parallel strategy; The third execution stage, the third execution stage including at least one third operator, where the third operator is an operator for which the data matrix of the operation meets the parallel requirement but the weight matrix of the operation does not meet the parallel requirement, and the third execution stage corresponds to the data parallel strategy; The fourth execution stage, the fourth execution stage including at least one fourth operator, where the fourth operator is an operator for which both the data matrix and the weight matrix of the operation meet the parallel requirement, and the fourth execution stage corresponds to the hybrid parallel strategy.
2. The method according to claim 1, wherein The parallel strategy corresponding to the execution stage is the conventional parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following: If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy, for any one of the artificial intelligence chips, retain the operation results of the artificial intelligence chip in the execution stage in the memory resources of the artificial intelligence chip itself; If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, divide the operation results of the artificial intelligence chip in the execution stage into a plurality of first data sub-matrices corresponding to each artificial intelligence chip respectively according to the number of artificial intelligence chips and with the row division dimension, and retain the first data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself; If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, divide the operation results of the artificial intelligence chip in the execution stage into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively according to the number of artificial intelligence chips, and retain the second data sub-matrix corresponding to the artificial intelligence chip itself in the memory resources of the artificial intelligence chip itself, where the division method of the second data sub-matrix includes: When the data matrix in the next execution stage is the left matrix of the multiplication operation, divide the operation results of the artificial intelligence chip in the execution stage into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively with the column division dimension; or when the data matrix in the next execution stage is the right matrix of the multiplication operation, divide the operation results of the artificial intelligence chip in the execution stage into a plurality of second data sub-matrices corresponding to each artificial intelligence chip respectively with the row division dimension.
3. The method according to claim 1, wherein The parallel strategy corresponding to the execution stage is the data parallel strategy. Based on the parallel strategy corresponding to the next execution stage, writing part or all of the operation results of each artificial intelligence chip in the execution stage directly into the memory resources of the target artificial intelligence chip includes at least one of the following: If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy or the tensor parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chip in the execution stage into the memory resources of each artificial intelligence chip respectively; If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, according to the number of artificial intelligence chips, with columns as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple third data sub-matrices corresponding to each artificial intelligence chip respectively, and write the multiple third data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves.
4. The method according to claim 1, characterized in that, The parallel strategy corresponding to the execution stage is the tensor parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip includes at least one of the following: If the parallel strategy corresponding to the next execution stage is the conventional parallel strategy, for any one of the artificial intelligence chips, write the operation results of the artificial intelligence chips in the execution stage into the memory resources of each artificial intelligence chip respectively; If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fourth data sub-matrices corresponding to each artificial intelligence chip respectively, and write the multiple fourth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips; If the parallel strategy corresponding to the next execution stage is the hybrid parallel strategy, for any one of the artificial intelligence chips, when the data matrix in the next execution stage is the left matrix of the multiplication operation, retain the operation results of the artificial intelligence chips in the execution stage in the memory resources of the artificial intelligence chips themselves; or, when the data matrix in the next execution stage is the right matrix of the multiplication operation, according to the number of artificial intelligence chips, with rows as the division dimension, divide the operation results of the artificial intelligence chips in the execution stage into multiple fifth data sub-matrices corresponding to each artificial intelligence chip respectively, and write the multiple fifth data sub-matrices into the memory resources of their respective corresponding artificial intelligence chips.
5. The method according to claim 1, characterized in that The parallel strategy corresponding to the execution stage is the hybrid parallel strategy. Based on the parallel strategy corresponding to the next execution stage, directly writing some or all of the operation results of each artificial intelligence chip in the execution stage into the memory resources corresponding to the target artificial intelligence chip includes: Perform a full reduction process on the operation results of each artificial intelligence chip in the execution stage, and perform at least one of the following on the full reduction process results: If the parallel strategy corresponding to the next execution stage is the tensor parallel strategy or the conventional parallel strategy, for any one of the artificial intelligence chips, store the all-reduction processing result in the memory resource of the artificial intelligence chip itself; If the parallel strategy corresponding to the next execution stage is the data parallel strategy, for any one of the artificial intelligence chips, divide the all-reduction processing result into multiple sixth data sub-matrices corresponding to each artificial intelligence chip in the dimension of row division according to the number of artificial intelligence chips, and store the sixth data sub-matrix corresponding to the artificial intelligence chip itself in the memory resource of the artificial intelligence chip itself.
6. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
System and method for executing automated machine learning solution, and electronic apparatus
WO2021052422A1
Data processing method and apparatus
WO2025030869A1