Tensor determination method and device based on expert parallelism and storage medium
Through expert parallel tensor determination method, the input tensor is split into sub-tenster parallel computing, which solves the problems of large memory usage and time-consuming data copying of deep learning network models, and improves the overall performance of the model.
Patent Information
- Application Number
- CN202510765512.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, deep learning network models have a great demand for video memory, and the data copying process between the CPU and the artificial intelligence chip is time-consuming, resulting in performance losses.
Using expert parallel tensor determination method, the input tensor is split into multiple sub-tensors, and distributed to multiple computing units in parallel for calculation. The first weight parameter of the expert sub-network is calculated, which reduces the memory space usage, and optimizes the weight parameters through load balancing and distributed optimizers.
It reduces video memory usage, reduces data copying between the CPU and artificial intelligence chip, improves model training and inference efficiency, and improves overall performance.
Smart Images

Figure CN120278237A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of artificial intelligence chips, and in particular, to a method, device, and storage medium for determining tensors based on expert parallelism. Background Art
[0002] Deep learning network models mainly include two development directions: horizontal and vertical. As the models develop deeper and deeper in both horizontal and vertical directions, the demand of the models for the video memory of artificial intelligence chips is also increasing.
[0003] To reduce the occupancy of video memory by the model, related technologies optimize the video memory by using the method of video memory offload, that is, temporarily store some video memory data in the Central Processing Unit (CPU) to reduce the occupancy of the video memory. Until this part of the video memory data is needed next time, it is copied from the CPU to the video memory again.
[0004] However, the copying of a large amount of video memory data between the CPU and the artificial intelligence chip is an extremely time-consuming process, which will cause a large performance loss to the entire model. Summary of the Invention
[0005] Embodiments of the present application provide a method, device, and storage medium for determining tensors based on expert parallelism, which are used to reduce the occupancy of video memory by the model and reduce the data copying between the CPU and the artificial intelligence chip, thereby improving the overall performance of the model.
[0006] On the one hand, embodiments of the present application provide a method for determining tensors based on expert parallelism, and the method includes: At the hidden layer dimension of the input tensor, split the input tensor into N sub-tensors, where N is greater than 1; Allocate the N sub-tensors to the expert sub-networks respectively deployed by N computing units; each expert sub-network includes corresponding first weight parameters, and the shape of each first weight parameter includes a target dimension associated with the hidden layer dimension; the length of the sub-tensor allocated to each expert sub-network in the hidden layer dimension is equal to the length of the first weight parameter of each expert sub-network in the target dimension; Through N expert sub-networks, calculate the N sub-tensors in parallel based on their respective first weight parameters to obtain corresponding sub-calculation results; Based on the obtained N sub-calculation results, obtain the output tensor.
[0007] On the one hand, embodiments of the present application provide a device for determining tensors based on expert parallelism, and the device includes: A splitting module, configured to split an input tensor into N sub-tensors in the hidden layer dimension of the input tensor, where N is greater than 1; An allocation module, configured to allocate the N sub-tensors to the expert sub-networks respectively deployed by N computing units; each expert sub-network includes corresponding first weight parameters, and the shape of each first weight parameter includes a target dimension associated with the hidden layer dimension; the length of the sub-tensor allocated to each expert sub-network in the hidden layer dimension is equal to the length of the first weight parameter of each expert sub-network in the target dimension; An execution module, configured to perform calculations on the N sub-tensors in parallel based on their respective first weight parameters through N expert sub-networks to obtain corresponding sub-calculation results; and obtain an output tensor based on the obtained N sub-calculation results.
[0008] Optionally, the N computing units are located on M artificial intelligence chips, where M is greater than 0 and less than or equal to N; The first weight parameters of the expert sub-networks deployed by each computing unit are stored in the video memory of the artificial intelligence chip where each computing unit is located.
[0009] Optionally, the execution module is further configured to: When the remaining storage space of the video memory does not meet the storage condition of the first weight parameters, transfer some of the data stored in the video memory to the central processing unit.
[0010] Optionally, the first weight parameters of the expert sub-network include: an upper projection layer parameter and a lower projection layer parameter; The target dimension in the shape of the upper projection layer parameter is: the row dimension; The target dimension in the shape of the lower projection layer parameter is: the column dimension.
[0011] Optionally, the execution module is further configured to: Update the length of the first weight parameter of each expert sub-network in the target dimension through a load balancing loss until the first weight parameters of the N expert sub-networks meet a preset balancing condition.
[0012] Optionally, the execution module is further configured to: Through each computing unit, respectively perform the following operations: Collect N first weight parameters to obtain a first global parameter; Split the first global parameter into N second weight parameters, and distribute the N second weight parameters to other computing units; the lengths of the N second weight parameters in the target dimension are the same; Copy one of the allocated second weight parameters to a distributed optimizer for parameter update to obtain a third weight parameter; Collect N third weight parameters, obtain a second global parameter, and obtain a fourth weight parameter from the second global parameter. The length of the fourth weight parameter in the target dimension is the same as that of the first weight parameter of the expert sub-network deployed by the computing unit.
[0013] Optionally, the execution module is further configured to: Before copying an assigned second weight parameter to a distributed optimizer for parameter update to obtain a third weight parameter, adjust the second weight parameter with the original precision to the second weight parameter with the target precision, where the target precision is greater than the original precision.
[0014] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned tensor determination method based on expert parallelism.
[0015] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned tensor determination method based on expert parallelism.
[0016] In the embodiment of the present application, in the hidden layer dimension of the input tensor, the input tensor is split into N sub-tensors; the N sub-tensors are distributed to the expert sub-networks respectively deployed by N computing units for parallel calculation. Since each expert sub-network calculates the assigned sub-tensor through the corresponding first weight parameter, the first weight parameter of the expert sub-network needs to be shape-matched with the assigned sub-tensor. Based on this, for the target dimension in the first weight parameter that is associated with the hidden layer dimension of the sub-tensor, the length of the target dimension is set to the length of the hidden layer dimension of the sub-tensor. In this way, while ensuring the accurate calculation of the sub-tensor, the length of the first weight parameter in the target dimension is reduced, that is, the video memory space occupied by the weight parameter of the expert sub-network itself is reduced.
[0017] Since the video memory space occupied by the weight parameter of the expert sub-network is reduced, the video memory data copy between the CPU and the artificial intelligence chip can be correspondingly reduced, thereby improving the efficiency of model training and inference, and further enhancing the overall performance of the model. In addition, since the total number of parameters of the weight parameter of the expert sub-network is reduced, the amount of data for data calculation and data communication is correspondingly reduced, thereby enhancing the performance of the entire mixture-of-experts network. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 A schematic structural diagram of an artificial intelligence chip provided by an embodiment of the present application; Figure 2 A schematic flow diagram of a tensor determination method based on expert parallelism provided by an embodiment of the present application; Figure 3 A schematic structural diagram of a hybrid expert model provided by an embodiment of the present application; Figure 4 A schematic flow diagram of a weight parameter update method provided by an embodiment of the present application Figure 1 ; Figure 5 A schematic flow diagram of a weight parameter update method provided by an embodiment of the present application Figure 2 ; Figure 6 A schematic structural diagram of a tensor determination device based on expert parallelism provided by an embodiment of the present application; Figure 7 A schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0021] Refer to Figure 1 , which is a structural diagram of an artificial intelligence chip applicable to an embodiment of the present application. The artificial intelligence chip 100 at least includes: a video memory 101 and a plurality of computing units 102. Among them, the computing unit 102 can be a Streaming Processing Cluster (SPC for short). The video memory 101 can be a High Bandwidth Memory (HBM for short) or other types of memories.
[0022] In an embodiment of the present application, the mixture of experts network includes multiple expert sub-networks, each expert sub-network is deployed in a computing unit 102, and the first weight parameters of each expert sub-network are stored in the video memory 101. The multiple expert sub-networks can be deployed in the computing units 102 included in one artificial intelligence chip 100, or can be deployed in the computing units 102 included in multiple artificial intelligence chips 100. In this regard, the present application does not make specific limitations.
[0023] When performing tensor determination based on expert parallelism, at the hidden layer dimension of the input tensor, the input tensor is split into N sub-tensors; then the N sub-tensors are allocated to the expert sub-networks respectively deployed in N computing units 102 for parallel computing. Since each expert sub-network calculates the allocated sub-tensor through the corresponding first weight parameter, therefore, the first weight parameter of the expert sub-network needs to be adapted to the allocated sub-tensor in terms of shape; based on this, for the target dimension associated with the hidden layer dimension of the sub-tensor in the first weight parameter, the length of the target dimension is set to the length of the hidden layer dimension of the sub-tensor.
[0024] Compared with the traditional solution (the length of the first weight parameter in the target dimension is: the length of the input tensor in the hidden layer dimension), the present application reduces the length of the first weight parameter in the target dimension, that is, reduces the video memory space occupied by the weight parameters of the expert sub-network itself. Then, when storing the first weight parameters of each expert sub-network through the video memory 101, the storage space occupied by the first weight parameters in the video memory 101 is greatly reduced, thereby reducing the storage pressure on the video memory 101.
[0025] Since the video memory space occupied by the weight parameters of the expert sub-network is reduced, the video memory data copy between the CPU and the artificial intelligence chip 100 can be correspondingly reduced, thereby improving the efficiency of model training and inference. In addition, since the total number of parameters of the weight parameters of the expert sub-network is reduced, therefore, the amount of data for the computing unit 102 to perform data calculation and data communication is correspondingly reduced, thereby improving the overall performance of the artificial intelligence chip 100.
[0026] In addition to including the above structures, the artificial intelligence chip 100 in the present application may further include other structures. In this regard, the present application does not make specific limitations.
[0027] The artificial intelligence chip 100 can be: a Graphics Processing Unit (GPU), a General-purpose computing on graphics processing units (GPGPU), a Domain Specific Architecture (DSA), etc.
[0028] Based on the following Figure 1 architecture diagram of the artificial intelligence chip, a process of providing a tensor determination method based on expert parallelism is specifically introduced. Refer to Figure 2 , this method is executed by a computer device, and the computer device includes Figure 1 the artificial intelligence chip shown in the figure. This method includes the following steps: Step 201, split the input tensor into N sub-tensors in the hidden layer dimension of the input tensor.
[0029] Specifically, N is greater than 1; the tensor determination method based on expert parallelism can be applied to various scenarios. For example, image processing scenarios, speech processing scenarios, text processing scenarios, etc. In different application scenarios, the physical meaning of the input tensor can be different.
[0030] For example, in the text processing scenario, the input tensor can be text data used in tasks such as text generation and text recognition.
[0031] For example, in the speech processing scenario, the input tensor can be speech data used in tasks such as speech enhancement, speech recognition, and speech synthesis.
[0032] For example, in the image processing scenario, the input tensor can be image data used in tasks such as image preprocessing, image segmentation, and object detection.
[0033] The shape of the input tensor includes one or more dimensions. When the input tensor is a two-dimensional tensor, the shape of the input tensor includes: a sequence dimension and a hidden layer dimension.
[0034] Taking the text processing scenario as an example, the original text "I love NLP " is obtained. The original text is split into a token sequence: "I", "love", "NLP", and this token sequence includes 3 tokens. An encoding operation is performed on each token in the token sequence to obtain a corresponding feature vector, and the length of this feature vector is 1024.
[0035] The token sequence and the feature vector corresponding to each token form the input tensor. Among them, the token sequence corresponds to the sequence dimension, and the length of the sequence dimension is 3; the feature vector corresponds to the hidden layer dimension, and the length of the hidden layer dimension is 1024.
[0036] In the sequence dimension and the hidden layer dimension, the sequence dimension is the row dimension, and the hidden layer dimension is the column dimension. Taking the text processing scenario as an example, the size of the input tensor corresponding to the original text "I love NLP " is: .
[0037] The central processing unit splits the input tensor into N sub-tensors at the hidden layer dimension of the input tensor. Among them, the lengths of the N sub-tensors in the hidden layer dimension can be the same or different; the lengths of the N sub-tensors in the sequence dimension are the same.
[0038] Step 202: Allocate the N sub-tensors to the expert sub-networks respectively deployed by N computing units.
[0039] Specifically, the mixture-of-experts network includes multiple expert sub-networks. In the embodiments of the present application, the mixture-of-experts network can be an independent model or a sub-network in a deep learning model.
[0040] For example, it is assumed that the deep learning model includes 16 layers in the vertical dimension. According to pipeline parallelism (PP for short), the deep learning model is split into 4 groups of networks, and each group of networks is deployed in an artificial intelligence chip, and each group of networks includes 4 layers.
[0041] Taking one group of networks as an example, see Figure 3 , the 4 layers in one group of networks are respectively: Self-Attention layer, Residual connection (Add)+Layer Normalization layer, Feed-Forward Network (FFN for short) layer, Residual connection+Layer Normalization layer. Among them, the Feed-Forward Network layer is the architecture of the mixture-of-experts network, including 4 expert sub-networks, namely FFN1, FFN2, FFN3, and FFN4, a router, and a merging unit. The router is used to allocate 4 sub-tensors to 4 expert sub-networks for parallel computing (i.e., expert parallelism), and the merging unit is used to merge the calculation results output by 4 expert sub-networks into a final result for output.
[0042] In addition, data parallelism (DP for short) and expert parallelism can share multiple artificial intelligence chips. Specifically, when data parallelism and expert parallelism are adopted in the deep learning model, expert parallelism reuses the artificial intelligence chips allocated to data parallelism. That is to say, first turn on data parallelism and turn off expert parallelism, that is, multiple artificial intelligence chips are used for data parallelism; when the model executes to the Feed-Forward Network layer, turn off data parallelism and turn on expert parallelism, that is, multiple artificial intelligence chips are used for expert parallelism.
[0043] It should be noted that in addition to data parallelism, deep learning models can also adopt tensor parallelism (abbreviated as TP); similarly, expert parallelism can reuse the artificial intelligence chips allocated to tensor parallelism. When data parallelism, tensor parallelism, and expert parallelism are simultaneously adopted in a deep learning model, multiple artificial intelligence chips are allocated to data parallelism and tensor parallelism, and expert parallelism can reuse the artificial intelligence chips allocated to any one of data parallelism and tensor parallelism. This can not only reduce the video memory occupancy of each artificial intelligence chip, but also improve the overall performance of the deep learning model.
[0044] In the embodiments of the present application, each expert sub-network is deployed in a computing unit, and each expert sub-network includes corresponding first weight parameters, and the first weight parameters are two-dimensional tensors. Since each expert sub-network calculates the allocated sub-tensors through the corresponding first weight parameters, the shape of each first weight parameter includes a target dimension associated with the hidden layer dimension; the length of the sub-tensors allocated to each expert sub-network in the hidden layer dimension is equal to the length of the first weight parameter of each expert sub-network in the target dimension. In addition to the target dimension, the shape of the first weight parameter also includes a reference dimension, and the length of the reference dimension is determined by the type of the expert sub-network.
[0045] For example, the mixture-of-experts network includes 4 expert sub-networks, and the size of the input tensor is: , where the length of the hidden layer dimension is 8192. The central processing unit splits the input tensor into 4 sub-tensors in the hidden layer dimension, and the size of each sub-tensor is: . The central processing unit allocates each sub-tensor to an expert sub-network deployed in a computing unit.
[0046] The first weight parameter of each expert sub-network is a two-dimensional tensor, and the size of the first weight parameter is: . When performing matrix multiplication with the sub-tensor as the left matrix and the first weight parameter as the right matrix, the length of the column dimension (i.e., the hidden layer dimension) of the sub-tensor is equal to the row dimension (i.e., the target dimension) of the first weight parameter, and the length is 2048.
[0047] In some embodiments, N computing units are located on M artificial intelligence chips, M is greater than 0 and less than or equal to N; the first weight parameters of the expert sub-networks deployed in each computing unit are stored in the video memory of the artificial intelligence chip where each computing unit is located.
[0048] Specifically, the number of expert sub-networks deployed on each artificial intelligence chip is the same. Corresponding storage spaces are allocated for the first weight parameters of each expert sub-network in the video memory.
[0049] For example, the mixture of experts network includes: 8 expert sub-networks, which are deployed on 4 artificial intelligence chips, and 2 expert sub-networks are deployed in each artificial intelligence chip. For each artificial intelligence chip, the two expert sub-networks are respectively deployed on two computing units, and the first weight parameters of each expert sub-network are stored in the video memory of the artificial intelligence chip.
[0050] Step 203: Through N expert sub-networks, calculate N sub-tensors in parallel based on their respective first weight parameters to obtain corresponding sub-computation results.
[0051] Step 204: Based on the obtained N sub-computation results, obtain an output tensor.
[0052] Specifically, in each computing unit, perform a matrix multiplication calculation on the input tensor and the first weight parameters of the expert sub-network to obtain corresponding sub-computation results, and the sub-computation results can be two-dimensional tensors. Then, splice the obtained N sub-computation results to obtain an output tensor, where the size of the output tensor is the same as the size of the input tensor.
[0053] In the embodiment of the present application, in the hidden layer dimension of the input tensor, the input tensor is split into N sub-tensors; the N sub-tensors are allocated to the expert sub-networks respectively deployed in N computing units for parallel calculation. Since each expert sub-network calculates the allocated sub-tensor through the corresponding first weight parameter, therefore, the first weight parameter of the expert sub-network needs to be adapted to the allocated sub-tensor in shape; based on this, for the target dimension associated with the hidden layer dimension of the sub-tensor in the first weight parameter, the length of this target dimension is set to the length of the hidden layer dimension of the sub-tensor. In this way, while ensuring the accurate calculation of the sub-tensor, the length of the first weight parameter in the target dimension is reduced, that is, the video memory space occupied by the weight parameter itself of the expert sub-network is reduced.
[0054] Since the video memory space occupied by the weight parameter of the expert sub-network is reduced, correspondingly, the video memory data copy between the CPU and the artificial intelligence chip can be reduced, thereby improving the efficiency of model training and inference, and further improving the overall performance of the model. In addition, since the total number of parameters of the weight parameter of the expert sub-network is reduced, therefore, the amount of data calculation and data communication is correspondingly reduced, thereby improving the performance of the entire mixture of experts network.
[0055] In some embodiments, the first weight parameter of the expert sub-network includes: an upper projection layer parameter and a lower projection layer parameter; the target dimension in the shape of the upper projection layer parameter is: the row dimension; the target dimension in the shape of the lower projection layer parameter is: the column dimension.
[0056] Specifically, both the upper projection layer parameters and the lower projection layer parameters are two-dimensional tensors. The input tensor is multiplied by the upper projection layer parameters through matrix multiplication to obtain the upper projection result, where the size of the upper projection result is larger than that of the input tensor. The upper projection result is multiplied by the lower projection layer parameters through matrix multiplication to obtain the sub-computation result, where the size of the sub-computation result is the same as that of the input tensor.
[0057] For example, it is assumed that the mixture-of-experts network includes 8 expert sub-networks, and the size of the input tensor is: , that is, the length of the sequence dimension is 2048, and the length of the hidden layer dimension is 8192.
[0058] In the traditional mixture-of-experts network, the input tensor is split into 8 sub-tensors in the sequence dimension, and the size of each obtained sub-tensor is: , . Then, the size of the upper projection layer parameters of the mixture-of-experts network is: , and the size of the lower projection layer parameters is: , where the parameter values "28672" and "14336" are determined by the type of the mixture-of-experts network. That is to say, for each expert sub-network, the upper projection layer parameters with a size of and the lower projection layer parameters with a size of need to be stored in the video memory, which greatly increases the video memory occupancy.
[0059] In the embodiment of the present application, the input tensor is split into 8 sub-tensors in the hidden layer dimension, and the size of each obtained sub-tensor is: , . Then, the size of the upper projection layer parameters of the mixture-of-experts network is: , and the size of the lower projection layer parameters is: , where "28672" and "14336" are determined by the type of the mixture-of-experts network. That is to say, for each expert sub-network, the upper projection layer parameters with a size of and the lower projection layer parameters with a size of need to be stored in the video memory.
[0060] Obviously, compared with the traditional splitting method, the method in the embodiment of the present application greatly reduces the sizes of the upper projection layer parameters and the lower projection layer parameters, thereby effectively reducing the video memory load of each artificial intelligence chip and reducing the parameter video memory pressure. Secondly, since the total number of parameters is reduced, the amount of computation corresponding to the parameters is also reduced, which will also improve the performance of the entire mixture-of-experts network structure to a certain extent.
[0061] In addition, in the current big data era, the input features may vary widely, including speech, text, images, coordinates, and so on. In order to input the rich and diverse features into the model to obtain an ideal result, in actual operation, each feature needs to be first mapped to a unique feature vector, and pruning is performed according to the similarity between the feature vectors and the correlation between the feature vectors and the model accuracy, so as to obtain a set of feature vectors with high accuracy and high performance for model training. Drawing on the above background, the hidden layer dimension of this application can be the dimension of the feature vector of the token sequence. The split based on the hidden layer dimension in this application can be understood as the split of large-scale independent feature vectors. At the same time, the selection of different expert sub-networks for different token sequences in the mixture-of-experts network is improved to the selection of different expert sub-networks for different feature vectors. Therefore, the solution of this application has strong theoretical applicability.
[0062] The technical solution of the embodiment of this application is applicable to both the model training stage and the model application stage. In the model training stage, it can reduce the video memory requirements and the peak video memory during the training process of the mixture-of-experts network, and at the same time, there is some performance improvement for the overall mixture-of-experts network.
[0063] During the model training process, not only the specific values of the first weight parameters are adjusted, but also the shapes of the first weight parameters are adjusted, that is, the lengths of the first weight parameters in the target dimension are adjusted.
[0064] In some implementations, at the initial stage of model training, the first weight parameters of each expert sub-network are in a non-equilibrium state, that is, the lengths of the first weight parameters of some expert sub-networks in the target dimension are very large, and the video memory space occupied by these weight parameters is very large; while the lengths of the first weight parameters of some other expert sub-networks in the target dimension are very small, and the video memory space occupied by these weight parameters is very small. This situation results in the underutilization of the video memory.
[0065] In view of this, this application updates the lengths of the first weight parameters of each expert sub-network in the target dimension through the load balancing loss until the first weight parameters of the N expert sub-networks meet the preset equilibrium condition.
[0066] Among them, the equilibrium condition can be that the lengths of the first weight parameters of multiple expert sub-networks in the target dimension are the same, or the difference in the lengths of the first weight parameters of any two expert sub-networks in the target dimension is within the preset threshold; of course, the equilibrium condition can also be in other forms, and this application does not make specific limitations in this regard.
[0067] Taking the upper projection layer parameters in the first weight parameters as an example, the mixture-of-experts network includes N expert sub-networks, where N is greater than 1. At the initial stage of training, the size of the upper projection layer parameters of each expert sub-network x is: , where, , the lengths of the upper projection layer parameters of N expert sub-networks in the target dimension are different and vary greatly. After multiple rounds of iterative training, the upper projection layer parameters of the N expert sub-networks meet the preset equilibrium condition, that is, when reaching the equilibrium state, the size of each upper projection layer parameter is: , that is, the lengths of the N upper projection layer parameters in the target dimension are the same.
[0068] The mixture-of-experts network is gradually transitioned from a non-equilibrium state to an equilibrium state through the load balancing loss, so that the video memory space occupied by the first weight parameters of each expert sub-network is gradually balanced, thereby improving the video memory utilization rate and also improving the overall performance of the mixture-of-experts network.
[0069] In some embodiments, at the initial stage of model training, the first weight parameters of each expert sub-network are in a non-equilibrium state, which may cause the weight parameters of some expert sub-networks to occupy too much video memory, and even the video memory may be difficult to meet the storage requirements.
[0070] Based on this, in the embodiments of the present application, when the remaining storage space of the video memory does not meet the storage condition of the first weight parameter, a part of the data stored in the video memory is transferred to the central processing unit.
[0071] In specific implementation, the storage condition may be that the storage space required by the first weight parameter is greater than the remaining storage space of the video memory; or the ratio between the storage space required by the first weight parameter and the remaining storage space of the video memory is greater than a preset ratio; of course, it may also be other set conditions. In this regard, the present application does not make specific limitations.
[0072] At the initial stage of training, a part of the data stored in the video memory is transferred to the central processing unit, and this part of the data may be intermediate results generated during the model training or inference process. In this way, more storage space can be vacated in the video memory to store the weight parameters of the expert sub-network to ensure the performance of model training. When the mixture-of-experts network transitions from a non-equilibrium state to an equilibrium state, the transfer of a part of the data stored in the video memory to the central processing unit is reduced, thereby reducing the video memory data copy between the CPU and the artificial intelligence chip, and further improving the overall performance of the model.
[0073] In some embodiments, during the model training process, a distributed optimizer is often used to update the parameters of the first weight parameters of each expert sub-network. In practical applications, the distributed optimizer is a static component, that is, it is used to update the parameters of weight parameters with a fixed shape.
[0074] However, in the implementation of this application, the shape of the first weight parameter needs to be adjusted during the model training process, that is, the length of the first weight parameter in the target dimension is adjusted. That is to say, during the training process, the shape of the first weight parameter is dynamically changing, which makes it difficult for the distributed optimizer to adapt to the shape of the first weight parameter that needs to be updated.
[0075] In view of this, in the embodiments of this application, refer to Figure 4 , and each computing unit respectively performs the following operations: Step 401, collect N first weight parameters to obtain the first global parameter.
[0076] Specifically, through the all-gather operation, receive the first weight parameters sent by the other N - 1 computing units, and then splice the first weight parameters of the local expert sub-network with the received N - 1 first weight parameters according to the first splicing order to obtain the first global parameter, where the first splicing order determines the position of each first weight parameter in the first global parameter.
[0077] Step 402, divide the first global parameter into N second weight parameters and distribute the N second weight parameters to other computing units.
[0078] Specifically, through the scatter operation, in the target dimension of the first global parameter, evenly divide the first global parameter into N second weight parameters and distribute the N second weight parameters to other computing units. The lengths of the N second weight parameters in the target dimension are the same; that is to say, the sizes of the N second weight parameters are the same.
[0079] Step 403, copy the allocated second weight parameter to the distributed optimizer for parameter update to obtain the third weight parameter.
[0080] Specifically, each computing unit corresponds to a distributed optimizer, and each distributed optimizer is used to perform parameter update on the second weight parameter allocated to the corresponding computing unit, where the shape of the second weight parameter is always fixed.
[0081] In some implementations, since the large model uses low-precision training and a high-precision parameter will be retained during the parameter update at the distributed optimizer, therefore, in the embodiments of this application, adjust the second weight parameter with the original precision to the second weight parameter with the target precision, where the target precision is greater than the original precision; then copy the second weight parameter with the target precision to the distributed optimizer for parameter update to obtain the third weight parameter, where the third weight parameter corresponds to the target precision.
[0082] That is to say, the dynamics of the low-precision weight parameters (i.e., the first weight parameters) are retained, and the static nature of the distributed optimizer during parameter update is maintained by means of the high-precision parameters (i.e., the second weight parameters). In this way, during the model training process, the shape of the first weight parameters can be adjusted, and at the same time, the parameter values of the first weight parameters can be indirectly updated with the help of the distributed optimizer.
[0083] Step 404: Collect N third weight parameters, obtain the second global parameter, and obtain the fourth weight parameter from the second global parameter.
[0084] Specifically, through all-to-all communication, receive the third weight parameters sent by the other N - 1 computing units; then splice the locally updated third weight parameters with the received N - 1 third weight parameters according to the second splicing order to obtain the second global parameter. The second splicing order determines the position of each third weight parameter in the second global parameter, where the second splicing order can be the same as the first splicing order described above.
[0085] The computing unit obtains the fourth weight parameter from the second global parameter according to the length of the first weight parameter of the expert subnetwork deployed locally in the target dimension and the first splicing order. The size of the fourth weight parameter is the same as that of the first weight parameter; the fourth weight parameter is actually the first weight parameter after parameter update.
[0086] At this time, the fourth weight parameter corresponds to the target precision, that is, corresponds to high precision; for the convenience of subsequent model training, the fourth weight parameter with the target precision is adjusted to the fourth weight parameter with the original precision.
[0087] For example, see Figure 5 , the mixture-of-experts network includes N expert subnetworks, and the N expert subnetworks are deployed on N computing units. The N computing units are respectively: . The size of the first weight parameter of the expert subnetwork deployed on EP1 is: , and the precision is BF16 (i.e., 16-bit floating point number); the size of the first weight parameter of the expert subnetwork deployed on EP2 is: , and the precision is BF16; and so on, the size of the first weight parameter of the expert subnetwork deployed on EP N is: , and the precision is BF16 (i.e., 16-bit floating point number), where, .
[0088] Collect and splice N first weight parameters through an all-gather operation to obtain the first global parameter, with a size of: , and the precision is BF16.
[0089] The first global parameter is evenly switched to N second weight parameters through a scattering operation, and the N second weight parameters are distributed to the other N - 1 computing units. The size of the second weight parameter assigned to each computing unit is: , with a precision of BF16.
[0090] Adjust the precision of the second weight parameter to FP32 (i.e., 32-bit floating point number), and then copy the second weight parameter to the distributed optimizer for parameter update to obtain the third weight parameter. The size of the third weight parameter is: , with a precision of FP32.
[0091] Through all-to-all communication, obtain the fourth weight parameter of the expert subnetwork deployed in each computing unit. Specifically, the size of the fourth weight parameter of the expert subnetwork deployed by EP1 is: , with a precision of BF16; the size of the fourth weight parameter of the expert subnetwork deployed by EP2 is: , with a precision of BF16; and so on. The size of the fourth weight parameter of the expert subnetwork deployed by EP N is: , with a precision of BF16.
[0092] In the embodiments of the present application, the first weight parameter with a dynamically changing shape is converted into a second weight parameter with a fixed shape through all-gather operation and scattering operation, and the second weight parameter is added to the distributed optimizer for parameter update to obtain the third weight parameter. Then, with the obtained multiple third weight parameters, the first weight parameter (i.e., the fourth weight parameter) after parameter update is restored, realizing the indirect update of the parameter value of the first weight parameter by means of the distributed optimizer, thereby greatly reducing the video memory space occupied by the first weight parameter while ensuring the effect of model training.
[0093] Based on the same technical concept, the embodiments of the present application provide a structural schematic diagram of a tensor determination device based on expert parallelism, as shown in Figure 6 . The tensor determination device 600 based on expert parallelism includes: A splitting module 601, configured to split the input tensor into N sub-tensors in the hidden layer dimension of the input tensor, where N is greater than 1; An allocation module 602, configured to allocate the N sub-tensors to the expert subnets respectively deployed in N computing units; each expert subnet includes a corresponding first weight parameter, and the shape of each first weight parameter includes a target dimension associated with the hidden layer dimension; the length of the sub-tensor allocated to each expert subnet in the hidden layer dimension is equal to the length of the first weight parameter of each expert subnet in the target dimension; An execution module 603 is configured to calculate the N sub-tensors in parallel based on respective first weight parameters through N expert sub-networks to obtain corresponding sub-calculation results; and obtain an output tensor based on the obtained N sub-calculation results.
[0094] Optionally, the N computing units are located on M artificial intelligence chips, where M is greater than 0 and less than or equal to N; The first weight parameters of the expert sub-networks deployed by each computing unit are stored in the video memory of the artificial intelligence chip where each computing unit is located.
[0095] Optionally, the execution module 603 is further configured to: When the remaining storage space of the video memory does not meet the storage condition of the first weight parameters, transfer some data stored in the video memory to the central processing unit.
[0096] Optionally, the first weight parameters of the expert sub-networks include: upper projection layer parameters and lower projection layer parameters; The target dimension in the shape of the upper projection layer parameters is: the row dimension; The target dimension in the shape of the lower projection layer parameters is: the column dimension.
[0097] Optionally, the execution module 603 is further configured to: Update the length of the first weight parameters of each expert sub-network in the target dimension through load balancing loss until the first weight parameters of the N expert sub-networks meet a preset balancing condition.
[0098] Optionally, the execution module 603 is further configured to: Perform the following operations respectively through each computing unit: Collect N first weight parameters to obtain a first global parameter; Divide the first global parameter into N second weight parameters and distribute the N second weight parameters to other computing units; the lengths of the N second weight parameters in the target dimension are the same; Copy an assigned second weight parameter to a distributed optimizer for parameter update to obtain a third weight parameter; Collect N third weight parameters to obtain a second global parameter, and obtain a fourth weight parameter from the second global parameter, where the length of the fourth weight parameter in the target dimension is the same as that of the first weight parameter of the expert sub-network deployed by the computing unit.
[0099] Optionally, the execution module 603 is further configured to: Before copying an assigned second weight parameter to a distributed optimizer for parameter update and obtaining a third weight parameter, adjust the second weight parameter with the original precision to the second weight parameter with a target precision, where the target precision is greater than the original precision.
[0100] In an embodiment of the present application, in the hidden layer dimension of an input tensor, the input tensor is split into N sub-tensors; the N sub-tensors are distributed to expert sub-networks respectively deployed on N computing units for parallel computing. Since each expert sub-network calculates the assigned sub-tensor through a corresponding first weight parameter, therefore, the first weight parameter of the expert sub-network needs to be shape-matched with the assigned sub-tensor; based on this, for the target dimension in the first weight parameter that is associated with the hidden layer dimension of the sub-tensor, set the length of the target dimension to the length of the hidden layer dimension of the sub-tensor. In this way, while ensuring the accurate calculation of the sub-tensor, the length of the first weight parameter in the target dimension is reduced, that is, the video memory space occupied by the weight parameter of the expert sub-network itself is reduced.
[0101] Since the video memory space occupied by the weight parameter of the expert sub-network is reduced, correspondingly, the video memory data copy between the CPU and the artificial intelligence chip can be reduced, thereby improving the efficiency of model training and inference, and further enhancing the overall performance of the model. In addition, since the total number of parameters of the weight parameter of the expert sub-network is reduced, therefore, the amount of data for data calculation and data communication is correspondingly reduced, thereby enhancing the performance of the entire mixture-of-experts network.
[0102] In an embodiment of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0103] Based on the same technical concept, an embodiment of the present application provides a computer device, such as Figure 7 shown, including at least one artificial intelligence chip 100 and a memory 701 connected to at least one artificial intelligence chip 100. In an embodiment of the present application, the specific connection medium between the artificial intelligence chip 100 and the memory 701 is not limited, Figure 7 taking the connection between the artificial intelligence chip 100 and the memory 701 through a bus as an example. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0104] In an embodiment of the present application, the memory 701 stores instructions executable by at least one artificial intelligence chip 100. By executing the instructions stored in the memory 701, the at least one artificial intelligence chip 100 can perform the steps of the above-mentioned tensor determination method based on expert parallelism.
[0105] Among them, the artificial intelligence chip 100 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 701 and calling the data stored in the memory 701, the tensor determination based on expert parallelism can be achieved. Optionally, the artificial intelligence chip 100 may include one or more processing units. The artificial intelligence chip 100 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the artificial intelligence chip 100. In some embodiments, the artificial intelligence chip 100 and the memory 701 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0106] The artificial intelligence chip 100 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0107] The memory 701, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 701 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical discs, and so on. The memory 701 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 701 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0108] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium that stores a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned tensor determination method based on expert parallelism.
[0109] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned tensor determination method based on expert parallelism.
[0110] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0111] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 a flow or flows and / or block Figure 1 or blocks.
[0112] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 a flow or flows and / or block Figure 1 or blocks.
[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart flow or flows and / or block or blocks. Figure 1 a flow or flows and / or block Figure 1 or blocks.
[0114] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0115] It is apparent that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for determining a tensor based on expert parallelism, characterized in that Including: At the hidden layer dimension of the input tensor, splitting the input tensor into N sub-tensors, where N is greater than 1; Allocating the N sub-tensors to the expert sub-networks respectively deployed by N computing units; each of the expert sub-networks includes corresponding first weight parameters, and the shape of each of the first weight parameters includes a target dimension associated with the hidden layer dimension; the length of the sub-tensor allocated to each expert sub-network in the hidden layer dimension is equal to the length of the first weight parameter of each expert sub-network in the target dimension; Through the N expert sub-networks, calculating the N sub-tensors in parallel based on their respective first weight parameters to obtain corresponding sub-calculation results; Based on the obtained N sub-calculation results, obtaining an output tensor.
2. The method according to claim 1, wherein The N computing units are located on M artificial intelligence chips, where M is greater than 0 and less than or equal to N; The first weight parameters of the expert sub-networks deployed by each computing unit are stored in the video memory of the artificial intelligence chip where each computing unit is located.
3. The method according to claim 2, characterized in that, It further includes: When the remaining storage space of the video memory does not meet the storage condition of the first weight parameters, transferring some of the data stored in the video memory to the central processing unit.
4. The method according to claim 1, wherein The first weight parameters of the expert sub-networks include: upper projection layer parameters and lower projection layer parameters; The target dimension in the shape of the upper projection layer parameters is: row dimension; The target dimension in the shape of the lower projection layer parameters is: column dimension.
5. The method according to claim 1, characterized in that, It further includes: Updating the length of the first weight parameters of each expert sub-network in the target dimension through load balancing loss until the first weight parameters of the N expert sub-networks meet a preset balancing condition.
6. The method according to any one of claims 1 to 5, characterized in that It further includes: Through each computing unit, respectively performing the following operations: Collecting N first weight parameters to obtain a first global parameter; Splitting the first global parameter into N second weight parameters and distributing the N second weight parameters to other computing units; the N second weight parameters have the same length in the target dimension; Copying one of the allocated second weight parameters to a distributed optimizer for parameter update to obtain a third weight parameter; Collecting N third weight parameters to obtain a second global parameter, and obtaining a fourth weight parameter from the second global parameter, where the fourth weight parameter has the same length as the first weight parameter of the expert sub-network deployed by the computing unit in the target dimension.
7. The method according to claim 6, wherein Before copying one of the allocated second weight parameters to a distributed optimizer for parameter update to obtain a third weight parameter, it further includes: Adjusting the one second weight parameter with the original precision to the one second weight parameter with the target precision, where the target precision is greater than the original precision.
8. A computer device, comprising a memory, an artificial intelligence chip, and a computer program stored in the memory and running on the artificial intelligence chip, characterized in that, When the artificial intelligence chip executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, It stores a computer program executed by a computer device. When the computer program runs on the computer device, it causes the computer device to execute the steps of the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to perform the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic parallelization method and device for hybrid expert model, equipment and medium
CN119806829A
Cited By
Optimization method and device of hybrid expert system, computer equipment and readable storage medium
CN120449952A
Tensor processing method and device based on hybrid expert network, and storage medium
CN121072622A
Data processing method and apparatus, electronic device, storage medium, and program product
CN122674764A