Optimization Method, Device, Equipment, Medium and Program of Mixture of Experts Model
By performing structured sparse conversion and data arrangement optimization on the sparse data structure of the hybrid expert model, combined with the sparse calculation unit of the computing chip, the sparse operator is generated, which solves the problem of insufficient performance improvement in sparse calculations of the hybrid expert model, and improves computing efficiency and resource utilization.
Patent Information
- Application Number
- CN202510308465.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In the prior art, the hybrid expert model lacks hardware-level instruction support when performing sparse calculations, resulting in the unstructured sparse calculation technology being unable to effectively improve performance, especially in the matrix multiplication calculation of double sparse mode data.
By performing structured sparse conversion and data arrangement optimization on the sparse data structure of the hybrid expert model, combined with the sparse calculation unit of the computing chip, a sparse operator is generated and the original operator is updated to achieve structured sparse optimization.
The calculation, bandwidth and storage resource overhead in the sparse matrix multiplication operation process are optimized, and the acceleration performance of sparse computing hardware is fully utilized.
Smart Images

Figure CN119808860B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to an optimization method, device, electronic device, storage medium and program for a mixture-of-experts model. Background Art
[0002] With the continuous development of LLMs (Large Language Model, large language model, hereinafter referred to as large model), both the scale of model parameters and the computing requirements have increased significantly. Currently, the model parameters of newly derived large models are as high as 400 billion, which is a substantial increase compared with the 110 million parameters of early models. This rapid development poses significant challenges to deploying these LLMs within existing artificial intelligence infrastructures, especially considering the limitations imposed by the speed of hardware progress.
[0003] For a brand-new LLM model constructed based on the mixture-of-experts (MoE) computing layer, due to its enhanced generalization ability and the ability to effectively manage multimodal tasks, it has been widely integrated into emerging large language models. This architectural innovation poses unique requirements for storage, bandwidth, and computing resources. Solving the problems of the growth of large language model scale and new architectures is crucial for their effective deployment on contemporary artificial intelligence accelerators.
[0004] In the process of implementing the present invention, the inventors found that for the challenges of the above-mentioned brand-new computing load, current cutting-edge work mainly focuses on performance optimization and acceleration based on unstructured sparse computing technology. However, unstructured sparse computing technology lacks hardware-level instruction support on artificial intelligence acceleration devices, such as GPUs (Graphics Processing Unit, graphics processors), for example, lacks dedicated sparse digital arithmetic unit (sparse ALU) support. Therefore, it cannot bring much performance improvement. Currently, the functions of hardware-level sparse digital arithmetic units mainly support uniformly scaled structured sparse technology, and do not implement matrix multiplication operators for double-sparse mode data (that is, both the left and right operands of the operator are sparse data). Therefore, directly using existing solutions cannot perform efficient matrix multiplication calculations on double-sparse mode data. Summary of the Invention
[0005] The embodiments of the present invention provide an optimization method, device, electronic device, storage medium and program for a mixture-of-experts model, which can achieve structured sparse optimization of the mixture-of-experts model, so as to give full play to the acceleration performance of the sparse computing hardware of the mixture-of-experts model when calculating sparse matrix multiplication, and thus greatly optimize the computing, bandwidth, and storage resource overheads in the sparse matrix multiplication operation process.
[0006] According to one aspect of the present invention, there is provided an optimization method for a mixture of experts model, including:
[0007] Performing a sparse conversion on various sparsified data structures of the target mixture of experts model to obtain a structured sparse data structure;
[0008] Optimizing the data arrangement of the structured sparse data structure;
[0009] Combining the sparse computing units of the computing chip to optimize the operators of the target mixture of experts model to obtain sparse operators;
[0010] Updating the original operators of the target mixture of experts model according to the sparse operators to obtain a structured sparse mixture of experts model.
[0011] According to another aspect of the present invention, there is provided an optimization device for a mixture of experts model, including:
[0012] A data structure sparse conversion module, configured to perform a sparse conversion on various sparsified data structures of the target mixture of experts model to obtain a structured sparse data structure;
[0013] A structured sparse data structure optimization module, configured to optimize the data arrangement of the structured sparse data structure;
[0014] A sparse operator optimization module, configured to combine the sparse computing units of the computing chip to optimize the operators of the target mixture of experts model to obtain sparse operators;
[0015] A sparse operator update module, configured to update the original operators of the target mixture of experts model according to the sparse operators to obtain a structured sparse mixture of experts model.
[0016] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the optimization method for the mixture of experts model according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to execute the optimization method for the mixture of experts model according to any embodiment of the present invention when executed.
[0021] According to another aspect of the present invention, there is also provided a computer program product, including a computer program which, when executed by a processor, implements the optimization method of the mixture-of-experts model according to any embodiment of the present invention.
[0022] In the embodiments of the present invention, sparse conversion is performed on various sparse data structures of the target mixture-of-experts model to obtain a structured sparse data structure, and data arrangement optimization is performed on the structured sparse data structure. Further, the operators of the target mixture-of-experts model are optimized in combination with the sparse computing units of the computing chip to obtain sparse operators, and then the original operators of the target mixture-of-experts model are updated according to the sparse operators to obtain a structured sparse mixture-of-experts model. The above technical solutions can solve the problem that it is difficult to effectively improve the performance when the existing mixture-of-experts model performs performance optimization and acceleration based on unstructured sparse computing technology, and can realize the structured sparse optimization of the mixture-of-experts model to give full play to the acceleration performance of the sparse computing hardware of the mixture-of-experts model when calculating sparse matrix multiplication, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the process of sparse matrix multiplication operation.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0025] Figure 1 is a flowchart of an optimization method for a mixture-of-experts model provided by an embodiment of the present invention;
[0026] Figure 2 is a schematic flowchart of a process for performing structural optimization processing on a mixture-of-experts model provided by an embodiment of the present invention;
[0027] Figure 3 is a flowchart of another optimization method for a mixture-of-experts model provided by an embodiment of the present invention;
[0028] Figure 4 is a schematic diagram of the control structure of a model weight matrix and a matching set of compression matrices applicable to an embodiment of the present invention;
[0029] Figure 5It is a comparison schematic diagram of the data rearrangement result required and matched by a sparse computing unit applicable to the embodiments of the present invention;
[0030] Figure 6 It is a comparison structural schematic diagram of an intermediate activation value sparse matrix and a matching selection array applicable to the embodiments of the present invention;
[0031] Figure 7 It is a flowchart of another optimization method for a mixture of experts model provided by the embodiments of the present invention;
[0032] Figure 8 It is a schematic flowchart for optimizing the original operator of a mixture of experts model provided by the embodiments of the present invention;
[0033] Figure 9 It is a flowchart of another optimization method for a mixture of experts model provided by the embodiments of the present invention;
[0034] Figure 10 It is an implementation block diagram of a data chunking and multi-level transfer process applicable to the embodiments of the present invention;
[0035] Figure 11 It is a schematic diagram of a register mapping relationship applicable to the embodiments of the present invention;
[0036] Figure 12 It is a structural schematic diagram of a sparse computing unit cooperating with an intermediate register to perform calculations applicable to the embodiments of the present invention;
[0037] Figure 13 It is a schematic diagram of an optimization device for a mixture of experts model provided by the embodiments of the present invention;
[0038] Figure 14 It is a structural schematic diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0039] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0041] Figure 1 is a flowchart of an optimization method for a mixture of experts model provided by an embodiment of the present invention. Figure 2 is a schematic flowchart of structurally optimizing a mixture of experts model provided by an embodiment of the present invention. This embodiment is applicable to the situation of structurally optimizing a mixture of experts model. This method can be executed by an optimization device of the mixture of experts model. The device can be implemented in a software and / or hardware manner and is generally integrated in an electronic device. The electronic device can be a terminal device or a server device, as long as it can execute the optimization method of the mixture of experts model. The specific device type of the electronic device is not limited in the embodiments of the present invention. Correspondingly, as Figure 1 shown, the method includes the following operations:
[0042] S110. Perform sparse conversion on the sparse data structures of the target mixture of experts model to obtain a structured sparse data structure.
[0043] Among them, the target mixture of experts model refers to a specific mixture of experts model used to be loaded into a computing chip to perform a model inference task. The mixture of experts model is an efficient neural network architecture that improves the overall performance by combining multiple specialized sub-models. The core idea of this model is to integrate different "expert" networks, and each expert plays a role in its proficient field. These experts can be small multi-layer perceptrons (MLPs) or more complex large language models (LLMs). The model usually includes a gating network that is responsible for determining which expert should be activated and participate in the calculation of the output when processing a specific input.
[0044] Among them, the structured sparse data structure can be a data structure in a structured sparse data format applicable to the target mixture of experts model.
[0045] The representation and arrangement optimization technology for sparse data is a technical solution that can effectively optimize the computational, bandwidth, and storage resource overheads during model training, inference, and fine-tuning. In the prior art, when a mixture-of-experts model performs model calculations, it mainly involves sparsely compressing the sparse weight data to be calculated and the input sparse intermediate activation data in a specified sparse representation format, and arranging them in a specified storage format in memory for subsequent calculations. In the process of implementing the present invention, the inventor found that the calculation between sparse intermediate activation data and sparse weight data belongs to double-sparse calculation. Currently, the mainstream sparse compression and arrangement optimization technologies lack an efficient representation of the double-sparse features in the calculation process of the mixture-of-experts model. Therefore, directly using the existing sparse compression and arrangement optimization solutions cannot fully optimize the calculation process of the mixture-of-experts model. Based on this, the embodiments of the present invention creatively propose an efficient data compression method for the sparsified model weight matrix in the mixture-of-experts model to effectively optimize the storage resource overhead of the mixture-of-experts model.
[0046] As Figure 2 shown, when performing sparse conversion on the various sparsified data structures of the target mixture-of-experts model, the model weights of the target mixture-of-experts model can be sparsely converted to obtain a structured sparse data structure of the model weights, or the sparse intermediate activation data input to the target mixture-of-experts model can be sparsely converted to obtain a structured sparse data structure of the activation values.
[0047] By separately converting the various sparsified model weight matrices in the target mixture-of-experts model in a specific format into a structured sparse data structure of the model weights and performing optimized storage in the computing chip used to implement the mixture-of-experts model calculations, the storage resource overhead of the computing chip can be effectively reduced. In addition, when using the model weights in the form of a structured sparse data structure for model calculations, the sparse computing hardware in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0048] S120. Optimize the data arrangement of the structured sparse data structure.
[0049] After completing the sparse conversion of the various sparsified data structures of the target mixture-of-experts model, the data arrangement of the structured sparse data structure can be further optimized to further improve the computing efficiency of the mixture-of-experts model.
[0050] In an alternative embodiment of the present invention, the optimization of data arrangement for the structured sparse data structure may include: before the target mixture-of-experts model runs, storing the model weight data after sparse conversion of the target mixture-of-experts model in a transposed manner; when the sparse operator of the target mixture-of-experts model performs data operations in shared memory, performing a data transpose operation on the intermediate activation values.
[0051] In traditional mixture-of-experts model calculations, the matrix multiplication calculation is: xW, where x is the input and W is the model weight. To meet the hardware requirements of the acceleration unit, the model calculation process will be reconstructed into . From the perspective of continuous memory access and the additional overhead of transposition, the embodiments of the present invention optimize and update the data arrangement. Specifically, before the target mixture-of-experts model runs, the model weight data after sparse conversion of the target mixture-of-experts model is stored in a transposed manner, that is, the transpose operation of the model weight W is performed during the offline model pruning process and no additional overhead is introduced during the running stage. To make full use of the bandwidth from global memory to shared memory, when the sparse operator of the target mixture-of-experts model performs data operations in shared memory, a data transpose operation is performed on the intermediate activation values, that is, the transpose of the input value x is completed at the shared memory level. Specifically, the data transpose of the last layer utilizes the fast instructions provided by the GPU hardware to complete.
[0052] S130. Optimize the operators of the target mixture-of-experts model in combination with the sparse computing unit of the computing chip to obtain sparse operators.
[0053] Among them, the sparse computing unit may be a hardware computing unit in the computing chip. The original mixture-of-experts model may be a mixture-of-experts model that needs to optimize the internal operators. The embodiments of the present invention do not limit the model type and model structure of the original mixture-of-experts model.
[0054] Optionally, the computing chip can be understood as an integrated circuit for implementing a set of computing tasks (such as Internet of Things control, high-performance computing, or mobile computing, etc.). The computing chip may be a general-purpose computing chip, such as a CPU (Central Processing Unit, central processor) or a GPU, or a heterogeneous chip combination of "CPU + GPU equipped with a sparse acceleration computing unit", or a dedicated computing chip, such as an NPU (Neural Processing Unit, neural network processor). The embodiments of the present invention do not limit this.
[0055] Among them, a sparse operator refers to an operator that only operates on some elements, such as a sparse matrix in matrix multiplication. During the calculation process, the processing of sparse operators usually involves how to effectively store and calculate sparse matrices, and how to optimize the calculation performance of sparse operators. Sparse operators are mainly applied to process sparse matrices, in which most elements are zero and only a small number of non-zero elements. Sparse operators can reduce the use of storage space through compression storage methods, accelerate operations such as sparse matrix multiplication using algorithms such as the fast Fourier transform, improve the memory usage efficiency using memory optimization techniques (such as cache optimization, memory alignment), and can also utilize parallel computing techniques (such as multi-threading or distributed computing) to accelerate the operations of large-scale sparse matrices.
[0056] The current mixture-of-experts computing technology mainly focuses on using sparse computing technology to solve the redundancy problem of parameters, while ignoring the sparse characteristics of input data. In addition, there is a lack of using hardware instruction support to provide structured sparse computing acceleration in mixture-of-experts computing.
[0057] In the embodiments of the present invention, in order to achieve the computing acceleration of adapting to a sparse acceleration hardware computing unit, fully considering the sparse characteristics of the operation data of the sparse operator, a data chunking method can be adopted to combine with the sparse computing unit of the computing chip to optimize the computing, bandwidth, and storage resource overhead during the operator execution process, and through compilation optimization means such as operator fusion, generate various types of sparse operators applicable to the target mixture-of-experts model.
[0058] S140. Update the original operator of the target mixture-of-experts model according to the sparse operator to obtain a structured sparse mixture-of-experts model.
[0059] Among them, the original operator can be the original operator type of the target mixture-of-experts model before structural optimization.
[0060] It can be understood that the original operator of the target mixture-of-experts model is usually a matrix multiplication operator. The matrix multiplication operator usually performs a dense matrix multiplication operation. Correspondingly, after generating various types of sparse operators applicable to the target mixture-of-experts model, the type and function of the original operator of the target mixture-of-experts model can be considered, and the original operator can be updated and replaced with an adapted sparse operator to achieve the operator optimization and structural update of the target mixture-of-experts model, and obtain a target mixture-of-experts model that can be accelerated by hardware instructions. The optimized and updated target mixture-of-experts model is a structured sparse mixture-of-experts model.
[0061] Correspondingly, the optimized target hybrid expert model can take the sparse matrix of the structured sparse data structure as input, optimize the data flow of the sparse operator in this format, and at the same time arrange the calculation process of the calculation kernel, and use the sparse computing units in the hardware-level computing chip to optimize and accelerate the sparse operators, and finally inject them into the hybrid expert calculation process, thereby accelerating the execution of the hybrid expert model.
[0062] It can be seen that the operator optimization method of the above hybrid expert model can realize structured sparse transformation suitable for hybrid expert calculation, and finally replace the original hybrid expert model with the converted model, use the same method as the original model calculation, and realize calculation acceleration by calling specific sparse acceleration units.
[0063] The embodiment of the present invention performs sparse conversion on the sparse data structures of the target hybrid expert model to obtain a structured sparse data structure, and optimizes the data arrangement of the structured sparse data structure. Furthermore, the operator of the target hybrid expert model is optimized in combination with the sparse computing unit of the computing chip to obtain a sparse operator, and then the original operator of the target hybrid expert model is updated according to the sparse operator to obtain a structured sparse hybrid expert model. The above technical scheme can solve the problem that the performance of the existing hybrid expert model is difficult to effectively improve when the performance is optimized and accelerated based on the unstructured sparse computing technology, and can realize the structured sparse optimization of the hybrid expert model to give full play to the acceleration performance of the sparse computing hardware of the hybrid expert model when calculating sparse matrix multiplication, thereby greatly optimizing the calculation, bandwidth and storage resource overhead in the sparse matrix multiplication process.
[0064] Figure 3 is a flow chart of another hybrid expert model optimization method provided by an embodiment of the present invention. This embodiment is specific based on the above embodiment. In this embodiment, multiple specific optional implementation methods for sparse conversion of various sparse data structures of the target hybrid expert model are provided. Figure 3 As shown, the method of this embodiment may include:
[0065] S210: Load the target hybrid expert model into a computing chip, and obtain a sparse model weight matrix of each item in the target hybrid expert model.
[0066] In order to facilitate the implementation of subsequent model calculations, in an embodiment of the present invention, the model weight matrix of each model layer in the target hybrid expert model is processed into a sparse model weight matrix of a specific structure.
[0067] Among them, the sparsified model weight matrix contains multiple compression units of a set size. Each of the compression units contains at least one valid row and at least one sparse row. Each of the valid rows contains at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse pattern.
[0068] Specifically, each model weight matrix in the target mixture-of-experts model contains an integer number (one or more) of compression units. That is, the size (number of rows * number of columns) of each model weight matrix is divisible by the size of the compression unit. A compression unit can be understood as a small matrix of a specific size, and this small matrix has at least two matrix rows. Among the above at least two matrix rows, there is at least one valid row and at least one sparse row. A valid row can be understood as a matrix row that contains at least one valid data (non-zero), and a sparse row refers to a matrix row where all data are 0.
[0069] Furthermore, each valid row in the compression unit contains an integer number (one or more) of structured sparse computing storage units. Among them, all the structured sparse computing storage units included in the model weight matrix are in the same sparse pattern. Generally speaking, for the convenience of calculation, the size of the structured sparse computing storage unit can be adapted to the computing scale of the hardware computing unit in the computing chip.
[0070] The sparse pattern can be understood as the ratio of valid data (non-zero value data) to all data. For example, the sparse pattern can be 2:4 or 4:8, etc. That is, for a 2:4 sparse pattern, the structured sparse computing storage unit contains a total of 4 data, and 2 of them are non-zero data. The arrangement positions of the above two non-zero data in the structured sparse computing storage unit are not restricted.
[0071] In the embodiments of the present invention, two methods can be adopted to make the model weight matrices in the target mixture-of-experts model in the above clearly defined special structure: one is to strongly constrain the model weight matrices in each model layer during the training stage of the target mixture-of-experts model, so that each model weight matrix meets the requirements of the above special structure; the other is to strongly constrain the model weight matrices in each model layer during the model fine-tuning stage after model pre-training, so that each model weight matrix meets the requirements of the above special structure.
[0072] In this embodiment, the reason for choosing to limit each model weight matrix in the target mixture-of-experts model to the above special structure is to improve the computing efficiency and better adapt to the specific hardware computing unit in the computing chip when performing calculations based on the target mixture-of-experts model later. The determination method of the above special structure will not be elaborated here.
[0073] S220. Generate a set of compression matrices corresponding to each of the model weight matrices respectively.
[0074] Among them, the set of compression matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the structured sparse computing storage unit.
[0075] As described above, in each model weight matrix, there are a large number of 0-valued data. If the above model weight matrices are directly stored in the computing chip according to the original size of the model weight matrix, a large amount of storage units will be wasted. In addition, directly performing subsequent model calculations based on the model weight matrix of the original size has low computing efficiency and cannot efficiently utilize the hardware computing units in the computing chip.
[0076] In view of this, in the embodiments of the present invention, for the model weight matrix with the above special structure, a novel and efficient data compression method is creatively proposed to solve the above technical problems.
[0077] In the embodiments of the present invention, through a specific data compression method, each model weight matrix can be tightly stored in the form of a set of compression matrices. Specifically, three compression matrices corresponding to the model weight matrix are stored in the set of compression matrices, namely the compressed data matrix, the index matrix, and the metadata matrix.
[0078] Among them, the compressed data matrix is used to store the compressed data of non-zero data in the model weight matrix, the index matrix is used to store the positions of valid rows in the model weight matrix in the corresponding compression unit, and the metadata matrix is used to store the positions of non-zero data in the structured sparse computing storage unit.
[0079] Obviously, through the above data compression method, the specific position of each non-zero data in the original model weight matrix can be determined through three small matrices. This data compression method can also effectively reduce the consumption of storage resources in the computing chip and facilitate the implementation of calculations based on the model weight matrix.
[0080] In an optional embodiment of the present invention, the sparse model weight matrix specifically includes a plurality of compression units of M*V size; each of the compression units specifically includes d valid rows, where 1≤d<M; and each of the structured sparse computing storage units is in the sparse mode of N:L, L is the total number of data included in the structured sparse computing storage unit, and N is the number of non-zero data included in the structured sparse computing storage unit.
[0081] In an alternative embodiment of the present invention, the generation of the set of compression matrices corresponding to each of the model weight matrices may include: obtaining a current model weight matrix of size m*k being processed currently, where m is an integer multiple of M and k is an integer multiple of V; constructing a first matrix of size (m / M*d)*(k / (L / N)), and filling the non-zero elements of the current model weight matrix into the first matrix row by row to obtain a compressed data matrix corresponding to the current model weight matrix; constructing a second matrix of size (m / M*d)*(k / V), and performing a filling process on the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix; constructing a third matrix of size (m / M*d)*(k / (L / N)), and performing a filling process on the third matrix according to the positions of each of the structured sparse computing and storage units in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing and storage unit to obtain a metadata matrix corresponding to the current model weight matrix.
[0082] For ease of explanation, Figure 4 shows a schematic diagram of a control structure of a model weight matrix applicable to various embodiments of the present invention and a matching set of compression matrices.
[0083] Specifically, as Figure 4 shown is a schematic diagram of generating a matching set of compression matrices for a model weight matrix with m = 4 and k = 16. In this model weight matrix, there are 4 compression units with M = 2 and V = 8. In each compression unit, there is d = 1 valid row, and each valid row contains 2 structured sparse computing and storage units, and each structured sparse computing and storage unit is in a sparse mode with N:L being 2:4. That is, the A, B, …, L filled in each matrix position of this model weight matrix represent non-zero data, and the blank positions represent data filled with 0 values.
[0084] When constructing the compressed data matrix for this model weight matrix, first construct a first matrix of (m / M*d)*(k / (L / N)) being 2*8, and then, fill the non-zero data A, B, …, L in the model weight matrix into this first matrix row by row to obtain a matching compressed data matrix.
[0085] Similarly, when constructing the index matrix for the model weight matrix, first construct a second matrix with dimensions (m / M*d) * (k / V) of 2 * 2. Then, based on the positions of each compression unit in the current model weight matrix and the positions of each valid row within its respective compression unit, fill the second matrix to obtain the index matrix corresponding to the current model weight matrix.
[0086] In an alternative embodiment of the present invention, the step of filling the second matrix based on the positions of each compression unit in the current model weight matrix and the positions of each valid row within its respective compression unit to obtain the index matrix corresponding to the current model weight matrix may include: sequentially traversing a current compression unit in the current model weight matrix, and based on the position of the current compression unit in the current model weight matrix, locating d vertical matrix positions in the second matrix that match the current compression unit; identifying the row positions of each valid row within the current compression unit, and filling the identified row positions into the d vertical matrix positions; and repeating the operation of sequentially traversing a current compression unit in the current model weight matrix until all compression units in the current model weight matrix have been processed, thereby obtaining the index matrix corresponding to the current model weight matrix.
[0087] As Figure 4 shown, first traverse a 2 * 8 compression unit 1 containing A, B, C, and D in the model weight matrix. Since d = 1, based on the upper left corner position of compression unit 1 in the model weight matrix, locate the corresponding 1 vertical matrix position in the upper left corner of the second matrix. Then, identify the row positions of the valid rows within compression unit 1, namely the rows where A, B, C, and D are located, i.e., row 0. Then, fill 0 into the matrix position in the upper left corner of the second matrix. And so on, until the row positions of each valid row in the four compression units in the model weight matrix are filled into the second matrix respectively to obtain the matching index matrix.
[0088] Similarly, when constructing the metadata matrix for the model weight matrix, first construct a third matrix with dimensions (m / M*d) * (k / (L / N)) of 2 * 8. Then, based on the positions of each structured sparse computing and storage unit in the current model weight matrix and the positions of each non-zero data within its respective structured sparse computing and storage unit, fill the third matrix to obtain the metadata matrix corresponding to the current model weight matrix.
[0089] In an alternative embodiment of the present invention, the step of filling the third matrix according to the positions of the structured sparse computing and storage units in the current model weight matrix and the positions of the non-zero data in their respective structured sparse computing and storage units to obtain a metadata matrix corresponding to the current model weight matrix may include: sequentially traversing a current structured sparse computing and storage unit in the current model weight matrix, and locating N horizontal matrix positions in the third matrix that match the current structured sparse computing and storage unit according to the position of the current structured sparse computing and storage unit in the current model weight matrix; correspondingly filling the column positions where each non-zero data is located in the current structured sparse computing and storage unit into the N horizontal matrix positions; and returning to perform the operation of sequentially traversing a current structured sparse computing and storage unit in the current model weight matrix until the processing of all the structured sparse computing and storage units in the current model weight matrix is completed, so as to obtain the metadata matrix corresponding to the current model weight matrix.
[0090] Continue as Figure 4 shown, each structured sparse computing and storage unit included in the model weight matrix can be sequentially traversed in a row-by-row traversal manner. For example, first, the structured sparse computing and storage unit 1 containing non-zero elements A and B is traversed. Since this structured sparse computing and storage unit 1 is located in the upper left corner of the model weight matrix, starting from the upper left corner of the third matrix, 2 horizontal matrix positions with N = 2 can be selected. Then, the column positions where A and B are located in the structured sparse computing and storage unit 1, A is in the 0th column and B is in the 2nd column, are correspondingly filled into the first two column positions of the first row of the third matrix. And so on. After the positions of the non-zero data in the 2×4 structured sparse computing and storage units in the above model weight matrix are correspondingly filled into the third matrix, a matching metadata matrix is obtained.
[0091] It can be understood that by constructing the above compression data matrix, index matrix, and metadata, the sparsified model weight matrix with the above specific structure can be uniquely determined. Furthermore, each sparsified model weight matrix included in the mixture-of-experts model can be efficiently compressed and stored, thereby effectively reducing the storage overhead of the mixture-of-experts model on the configured computing chip. In addition, when the mixture-of-experts model deployed in the computing chip performs calculations based on the above compression-form model weight matrix, the hardware computing advantages of each sparse computing unit in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0092] S230. Store each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compression matrices to obtain a structured sparse data structure of the model weight data.
[0093] Specifically, when it is necessary to store the sparsified model weight matrices in the target mixture-of-experts model in the storage space (typically, global memory) of the computing chip, each set of compressed matrices obtained after compression is used to replace each sparsified model weight matrix for storage to achieve storage optimization.
[0094] The technical solution of the embodiment of the present invention, by respectively converting each sparsified model weight matrix in the target mixture-of-experts model with a specific format into a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing storage unit, and then performing optimized storage in the computing chip for implementing the calculation of the mixture-of-experts model, can effectively reduce the storage resource overhead of the computing chip. In addition, when using the model weight matrix in the above data compression form for model calculation, the sparse computing hardware in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0095] In an alternative embodiment of the present invention, after respectively storing each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compressed matrices, it may further include: obtaining a sparse computing unit requirement file matching the computing chip; wherein, the sparse computing unit requirement file defines the data access positions of each thread executed in the sparse operator to each model weight matrix in the target mixture-of-experts model during sparse computing; rearranging the data arrangement modes of the compressed data matrices, index matrices, and metadata matrices in each set of compressed matrices according to the sparse computing unit requirement file; and storing the rearranged sets of compressed matrices in the set memory of the computing chip to achieve continuous storage of data with continuous access characteristics.
[0096] Generally speaking, when a computing chip (e.g., GPU) loads data, it adopts the technology of Memory Access Coalescing. Its core idea is to merge the memory accesses of multiple threads into a larger memory access operation, which can reduce access latency and improve bandwidth utilization. At the same time, its data loading also has the following characteristics:
[0097] 1. Aligned access: Memory access in GPUs is usually aligned in a certain stride, typically 32 bytes, 64 bytes or larger. If multiple threads operate on consecutive memory addresses, these accesses can usually be merged into a single memory access operation, thereby reducing the memory access overhead. 2. Coalesced memory access: For example, if multiple threads access adjacent addresses in memory, the GPU can merge these accesses into a single memory request, avoiding multiple memory accesses, thereby reducing latency and bandwidth waste.
[0098] Based on the above data loading idea, in each embodiment of the present invention, according to the data reading requirements of the sparse computing units in the computing chips used for implementation of computations, the data arrangement modes of the compressed data matrices, index matrices, and metadata matrices in each set of compressed matrices are rearranged to further improve the data reading efficiency during subsequent implementation of computations.
[0099] Specifically, by parsing the sparse computing unit requirement file provided with the computing chip during factory shipment, the data reading requirements of the sparse computing units can be obtained. Specifically, in this sparse computing unit requirement file, the data access positions of each thread in the sparse computing unit for each model weight matrix in the target mixture-of-experts model are recorded.
[0100] Among them, the sparse computing unit can be understood as a special hardware circuit in the computing chip used for performing matrix multiplication calculations on sparse matrices. Such special hardware circuits can accelerate sparse matrix multiplication. Generally, a computing chip contains multiple sparse computing units. Inside each sparse computing unit, one or more threads can be enabled to perform corresponding sparse computing tasks.
[0101] By way of example and not limitation, in Figure 5 a comparison schematic diagram of a sparse computing unit requirement and the corresponding data rearrangement result applicable to the embodiments of the present invention is shown.
[0102] As Figure 5 shown, in the left table in Figure 5 , the access positions of each thread adapted in each sparse computing unit for the matrix elements in a specific matrix (for example, the compressed data matrix) are recorded. Among them, T0, T1, …, T 29 respectively represent threads, and different threads belong to different sparse computing units. T0{3...0} in the 0th row and the 0...3rd columns in the left table represents that the data in the 0...3rd columns of the 0th row in this compressed data matrix is loaded into the {3...0} data position in the target vector used by the T0 thread for computation. Furthermore, the T0 thread needs to read the data in the 0th row and the 0...3rd columns of this compressed data matrix from the storage space where the compressed data matrix is located and store it in the matching position in the target vector.
[0103] Similarly, T0{19...16} in columns 0...3 of row 8 in the left table represents that the data in columns 0...3 of row 8 in the compressed data matrix is loaded into the data position {19...16} in the target vector used by thread T0 for calculation. Furthermore, thread T0 needs to read the data in columns 0...3 of row 8 in the compressed data matrix from the storage space where the compressed data matrix is located and store it at the matching position in the target vector.
[0104] It can be understood that considering the calculation characteristics of the actual sparse matrix, each thread does not read data from consecutive positions in a matrix. This data reading method has low efficiency and poor bandwidth utilization, which will reduce the calculation efficiency of the double sparse matrix to a certain extent. Based on this, in various embodiments of the present invention, according to the data reading characteristics of each thread in the sparse calculation unit, the data arrangement methods of the compressed data matrix, index matrix, and metadata matrix in each compressed matrix set are rearranged.
[0105] Specifically, as shown in the left and right tables in Figure 5 All 32 calculation data required by thread T0 are located in columns 0...15 of row 0 and columns 0...15 of row 8 in the compressed data matrix. However, according to the conventional data storage method, these 32 data are not consecutive. In other words, a thread needs to perform two memory transfers to obtain the data required by the sparse calculation unit, which wastes the memory bandwidth of the calculation chip.
[0106] To improve memory utilization, as can be seen from the right table in Figure 5 the 32 non-consecutively stored data mentioned above can be continuously arranged in a new matrix. Through such optimization, each thread needs to load 32 consecutive bits of data, and the complete data loading operation can be performed through one memory transfer. At this time, thread T0 can perform continuous data reading on the data that needs to be continuously accessed, greatly reducing the memory access overhead and data reading latency, and effectively improving the bandwidth utilization.
[0107] S240. Optimize the operators of the target mixture-of-experts model in combination with the sparse calculation unit of the computing chip to obtain sparse operators.
[0108] S250. Update the original operators of the target mixture-of-experts model according to the sparse operators to obtain a structured sparse mixture-of-experts model.
[0109] S260. Update the original operators of the target mixture-of-experts model according to the sparse operators to obtain a structured sparse mixture-of-experts model.
[0110] In an alternative embodiment of the present invention, after storing the model weight matrices with various sparsifications in the target mixture-of-experts model as a set of matching compressed matrices respectively, it may further include: invoking each of the sparse operators on the computing chip to perform matching sparse computations based on the target mixture-of-experts model with storage optimization.
[0111] When the sparse conversion, data layout optimization, operator fusion, and operator injection processing of the target mixture-of-experts model are completed to obtain a structured sparse mixture-of-experts model. Further, data can be input to the optimized target mixture-of-experts model. The optimized target mixture-of-experts model can invoke each of the sparse operators on the computing chip to perform matching sparse computation processes on the input data.
[0112] In an alternative embodiment of the present invention, before invoking each of the sparse operators on the computing chip to perform matching sparse computations based on the target mixture-of-experts model with storage optimization, it may further include: identifying each mixture model layer included in the target mixture-of-experts model and obtaining an expert assignment pattern corresponding to each mixture model layer; wherein the expert assignment pattern is used to indicate the row position where the valid rows of the intermediate activation values assigned to its own mixture model layer are located during the implementation of the computation; generating a selection array corresponding to each mixture model layer according to the expert assignment pattern; and associatively storing the selection arrays of each mixture model layer with the set of compressed matrices of each model weight matrix for combining the selection array with the sparse matrix or dense matrix of the intermediate activation values to be processed during the implementation of the computation and performing sparse computations with the matching set of compressed matrices.
[0113] As described above, the characteristic of the mixture-of-experts model is that it is necessary to determine which expert should be activated and participate in the computation of the output when processing a specific input. That is, when performing computations based on the target mixture-of-experts model, for any mixture model layer, its intermediate activation values are assigned to different experts for processing by the intermediate routing layer. Optionally, the intermediate activation values may include two forms: a sparse matrix of intermediate activation values and a dense matrix of intermediate activation values.
[0114] The inventor found through research that when using the mixture-of-experts model to calculate tasks (typically, model inference tasks), the matrix multiplication calculation between the model weight matrix of each model layer and the sparse matrix of intermediate activation values input to the mixture model layer in the mixture-of-experts model belongs to double sparse matrix multiplication. The matrix multiplication calculation between the model weight matrix of each model layer and the dense matrix of intermediate activation values input to the mixture model layer in the mixture-of-experts model belongs to sparse matrix multiplication. Furthermore, the methods of the embodiments of the present invention can be applied to the model calculation and model optimization scenarios of the mixture-of-experts model.
[0115] Optionally, for any expert, the intermediate activation values that are not assigned to it for processing can thus become structured sparse values. That is, the input to each mixture model layer can be a sparse matrix of intermediate activation values in the form of a sparse matrix, and one or more matrix rows in this sparse matrix of intermediate activation values are sparse rows (all zeros). Optionally, the input to each mixture model layer can also be a dense matrix of intermediate activation values in the form of a dense matrix, and all matrix elements in this dense matrix of intermediate activation values are valid data (all non-zeros).
[0116] In the traditional processing method, for each expert in each mixture model layer, the data assigned from the original sparse matrix or dense matrix of intermediate activation values (the data in non-sparse rows) is copied into new data, which will bring additional computational, memory, and bandwidth requirements. Based on this, each embodiment of the present invention further proposes a structured sparse representation method based on column vectors to avoid the additional data copying process, thereby greatly reducing the runtime overhead and improving the model running efficiency.
[0117] That is, directly store the sparse matrix or dense matrix of intermediate activation values for each mixture model layer respectively, and avoid the additional cost brought by copying to obtain new data.
[0118] In the embodiments of the present invention, since the sparse matrix or dense matrix of intermediate activation values is directly stored corresponding to each mixture model layer respectively, therefore, in order to meet the computational requirements of the sparse computing unit for the sparse matrix or dense matrix of intermediate activation values, it is necessary to synchronously store a selection array for the sparse matrix or dense matrix of intermediate activation values.
[0119] Among them, the selection array is used to describe the positions of valid rows in the sparse matrix or dense matrix of intermediate activation values.
[0120] As an example but not a limitation, in Figure 6 shows a schematic diagram of the control structure of a sparse matrix of intermediate activation values applicable to the embodiments of the present invention and the matching selection array. In Figure 6 the left matrix shows the specific form of a sparse matrix of intermediate activation values adapted to a mixture model layer. In this sparse matrix of intermediate activation values, it is set that the 0th row, the 3rd row, and the 5th row are valid rows, and the remaining rows are all-zero sparse rows. At this time, for the expert allocation mode of this mixture model layer, the intermediate activation values output by the 0th, 3rd, and 5th experts. Furthermore, a selection array in the form of {0, 3, 5} can be constructed to describe the positions of the valid rows in this sparse matrix of intermediate activation values. Assuming that the matrix corresponding to the intermediate activation values is a dense matrix, and this dense matrix includes the valid data in the 1st - 3rd rows and does not include sparse rows, then a selection array in the form of {1, 2, 3} can be constructed and associated with this second association matrix for storage.
[0121] In the technical solution of the embodiment of the present invention, by respectively converting the sparse model weight matrices in specific formats in the target mixture-of-experts model into a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing and storage unit, and then performing optimized storage in the computing chip for implementing the mixture-of-experts model calculation, the storage resource overhead of the computing chip can be effectively reduced. In addition, when using the model weight matrix in the above data compression form for model calculation, the sparse computing hardware in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0122] In the embodiment of the present invention, since the target mixture-of-experts model needs to perform matrix multiplication on the model weight matrix of each model layer and the input activation value matrix input to the model layer during model inference, sparse matrix multiplication is involved in this calculation process. In each embodiment of the present invention, by storing the model weight matrix in a compressed manner as a set of compressed matrices and synchronously storing the selection array for each mixture model layer, the sparse computing unit in the computing chip can be fully utilized to improve the computing efficiency.
[0123] Figure 7 is a flowchart of another optimization method for the mixture-of-experts model provided by the embodiment of the present invention. Figure 8 is a schematic flowchart for optimizing the original operator of the mixture-of-experts model provided by the embodiment of the present invention. This embodiment is specific based on the above embodiment. In this embodiment, various specific and optional implementation manners for performing sparse conversion on the sparse data structures of the target mixture-of-experts model are given. Correspondingly, as Figure 7 and Figure 8 shown, the method of this embodiment may include:
[0124] S310. Perform sparse conversion on the sparse data structures of the target mixture-of-experts model to obtain a structured sparse data structure, and optimize the data arrangement of the structured sparse data structure.
[0125] S320. Generate a first sparse operator, a second sparse operator, and a third sparse operator applicable to the target mixture-of-experts model by combining the sparse computing unit of the computing chip in a data block manner.
[0126] Among them, the first sparse operator may be an operator for outputting the model calculation result in the mixture-of-experts model. The second sparse operator may be an operator capable of outputting the output value of the activation function. The third sparse operator may be an operator capable of outputting the output value of the dot product function.
[0127] In an alternative embodiment of the present invention, if the sparse operator includes a first sparse operator, the optimization of the operator of the target mixture-of-experts model by the sparse computing unit of the combined computing chip to obtain a sparse operator may include: locating a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip according to the first sparse matrix multiplication requirement; wherein the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; transporting the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data-block form level by level; calculating the multiplication result of the first sparse matrix and the second associated matrix step by step by the sparse computing unit of the computing chip according to the data loaded in batches in the hardware registers; and processing the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the broadcast multiplication function of the target mixture-of-experts model to obtain the first sparse operator.
[0128] In an alternative embodiment of the present invention, if the sparse operator includes a second sparse operator, the optimization of the operator of the target mixture-of-experts model by the sparse computing unit of the combined computing chip to obtain a sparse operator may include: locating a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip according to the second sparse matrix multiplication requirement; wherein the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; transporting the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data-block form level by level, and rearranging or deforming the data in the original global memory during the transportation process; calculating the multiplication result of the first sparse matrix and the second associated matrix step by step by the sparse computing unit of the computing chip according to the data loaded in batches in the hardware registers; and processing the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the activation function of the target mixture-of-experts model to obtain the second sparse operator.
[0129] In an alternative embodiment of the present invention, if the sparse operator includes a third sparse operator, the optimization of the operator of the target mixture-of-experts model by the sparse computing unit of the combined computing chip to obtain a sparse operator may include: locating, according to the third sparse matrix multiplication requirement, a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; transporting the set of compressed matrices and the second associated matrix in the form of data blocks from the global memory to the hardware registers of the computing chip step by step, and rearranging or transforming the data in the original global memory during the transportation process; calculating, by the sparse computing unit of the computing chip, the multiplication result of the first sparse matrix and the second associated matrix step by step according to the data loaded in batches in the hardware registers; and processing the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the dot product function of the target mixture-of-experts model to obtain the third sparse operator.
[0130] Among them, the first sparse matrix may be the left operand of the double matrix multiplication and is a sparse matrix. In the embodiments of the present invention, the first sparse matrix may be the sparsified model weight matrix obtained by performing sparse conversion processing on the above embodiments, that is, the structured sparse data structure of the model weight data. The second associated matrix may be the right operand of the double matrix multiplication and may be a sparse matrix or a dense matrix composed of valid data in the sparse matrix. Optionally, the second associated matrix may be the intermediate activation value sparse matrix input to the mixture model layer in the target mixture-of-experts model, and the dense matrix corresponding to the second sparse matrix may be the dense matrix of the intermediate activation value input to the mixture model layer in the target mixture-of-experts model. The matrix multiplication between the first sparse matrix and the second associated matrix may be referred to as sparse matrix multiplication. When the second associated matrix also uses a sparse matrix, the matrix multiplication between the first sparse matrix and the second associated matrix may be referred to as double sparse matrix multiplication. In the following text, regardless of which type of matrix the second associated matrix uses, the matrix multiplication between the first sparse matrix and the second associated matrix is simply referred to as sparse matrix multiplication.
[0131] In an embodiment of the present invention, the above-mentioned method for accelerating the multiplication of sparse matrices can be encapsulated in the form of an operator interface. For example, multiple types of sparse operators for implementing matrix multiplication between a first sparse matrix and a second associated matrix are specifically constructed, such as a first sparse operator, a second sparse operator, or a third sparse operator, etc. Furthermore, by calling the operator interface of the sparse operator, the method for accelerating the multiplication between the first sparse matrix and the second associated matrix can be triggered, and according to the fusion requirements of the first sparse operator and other functional functions, additional processing logic can be configured for the multiplication result between the first sparse matrix and the second associated matrix, thereby generating multiple types of sparse operators, such as a first sparse operator, a second sparse operator, or a third sparse operator, etc.
[0132] Specifically, sparse matrix multiplication can be understood as that at least one type of operand among the left operand and the right operand for performing matrix multiplication is a sparse matrix. Correspondingly, when a first sparse operator is generated upon detecting an interface call request for the first sparse operator, it can be determined that a first sparse matrix multiplication requirement is detected; when a second sparse operator is generated upon detecting an interface call request for the second sparse operator, it can be determined that a second sparse matrix multiplication requirement is detected; when a third sparse operator is generated upon detecting an interface call request for the third sparse operator, it can be determined that a third sparse matrix multiplication requirement is detected. Furthermore, the identification information of the left operand (hereinafter referred to as the first sparse matrix) and the right operand (hereinafter referred to as the second associated matrix) that need to perform the multiplication of the sparse matrix can be obtained from the corresponding interface call request. Based on the above identification information, the compressed matrix set and the second associated matrix that match the first sparse matrix can be located in the global memory of the computing chip.
[0133] It can be understood that the first sparse matrix and the second associated matrix are used to perform multiplication calculations in the computing chip. Furthermore, the first sparse matrix and the second associated matrix that need to be calculated are required to be pre-loaded into the computing chip in advance. To improve the computing speed of the computing chip, the computing data is initially stored in the global memory of the computing chip, and subsequently, through a step-by-step transfer method, the computing data can be transferred in blocks to the hardware register adapted to the hardware computing unit, and the hardware computing unit calculates the final computing result based on the data blocks stored in the hardware register.
[0134] In an embodiment of the present invention, in order to achieve the final multiplication acceleration of the sparse operator, a special limitation is imposed on the sparse format of the first sparse matrix as the left operand. At the same time, the storage method of the first sparse matrix in the global memory of the computing chip is also optimized accordingly.
[0135] Among them, the first sparse matrix includes multiple compression units of a set size. Each compression unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the compression matrix set includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in their respective compression units, and a metadata matrix for storing the positions of non-zero data in their respective structured sparse computing storage units.
[0136] Specifically, the first sparse matrix includes an integer number (one or more) of compression units. That is, the size (number of rows * number of columns) of the first sparse matrix can be divided evenly by the size of the compression unit. The compression unit can be understood as a small matrix of a specific size, and this small matrix has at least two matrix rows. Among the above at least two matrix rows, there is at least one valid row and at least one sparse row. A valid row can be understood as a matrix row that includes at least one valid data (non-zero), and a sparse row refers to a matrix row where all data are 0.
[0137] Furthermore, each valid row in the compression unit includes an integer number (one or more) of structured sparse computing storage units. Among them, all the structured sparse computing storage units included in the first sparse matrix are in the same sparse mode. Generally speaking, for the convenience of calculation, the size of the structured sparse computing storage unit can be adapted to the computing scale of the hardware computing unit in the computing chip.
[0138] It can be understood that in the first sparse matrix, there are a large number of 0-valued data. If the first sparse matrix is directly stored in the global memory of the computing chip according to its original size, it will cause a waste of a large number of storage units. In addition, directly performing subsequent sparse matrix multiplication based on the first sparse matrix of the original size has low computing efficiency and cannot efficiently use the hardware computing unit in the computing chip, that is, the sparse computing unit.
[0139] In view of this, in the embodiments of the present invention, for the left operand (the first sparse matrix) of the sparse multiplication calculation with the above special structure, a novel and efficient data compression method is creatively proposed to solve the above technical problems. In the embodiments of the present invention, through a specific data compression method, the first sparse matrix can be tightly stored in the form of a set of compressed matrices. Specifically, three compressed matrices corresponding to the first sparse matrix are stored in the set of compressed matrices, namely, a compressed data matrix, an index matrix, and a metadata matrix. Among them, the compressed data matrix is used to store the compressed data of the non-zero data in the first sparse matrix, the index matrix is used to store the positions of the valid rows in the first sparse matrix in the corresponding compression unit, and the metadata matrix is used to store the positions of the non-zero data in the corresponding structured sparse calculation storage unit.
[0140] Obviously, through the above data compression method, the specific positions of each non-zero data in the original first sparse matrix can be determined through three small matrices. This data compression method can also effectively reduce the consumption of storage resources in the computing chip and facilitate the implementation of calculations based on the first sparse matrix.
[0141] In the embodiments of the present invention, the second correlation matrix may not be compressed and stored. However, considering the actual application requirements of sparse matrix multiplication, for example, when applied in a mixture of experts model, the second correlation matrix generally also has a specific data structure. Typically, when the second correlation matrix is a second sparse matrix, one or more matrix rows in the second sparse matrix are sparse rows (all 0). When the second correlation matrix is the dense matrix corresponding to the second sparse matrix, the data in the dense matrix corresponding to the second sparse matrix are the valid data in the second sparse matrix. Exemplarily, assuming the second sparse matrix is [1, 3, 0, 5, 0, 0], the dense matrix corresponding to the second sparse matrix is [1, 3, 5].
[0142] Based on this, in the global memory of the computing chip, in addition to storing the second correlation matrix completely, a selection array matching the second correlation matrix is further stored to describe the positions of the valid rows in the second correlation matrix. Correspondingly, when locating the second correlation matrix in the global memory of the computing chip, the selection array matching the second correlation matrix can be located synchronously.
[0143] In an optional implementation manner of the embodiments of the present invention, the computing chip includes a three-level storage architecture of global memory -> shared memory -> hardware register. Among them, the above three-level storage is physically closer to the hardware computing unit used to implement calculations in the computing chip, especially the hardware register is closely arranged with the hardware computing unit.
[0144] Generally speaking, the closer the storage device is to the computing unit, the faster its data read / write speed, but its storage capacity is also smaller. Therefore, for a matrix multiplication of C(M×N)=A(M×K)×B(K×N), the complete data of A and B cannot be fully stored in the shared memory or hardware registers. Instead, it is often necessary to transfer the data in chunks through multiple levels of transfer to the hardware registers, so that the hardware computing units in the computing chip can calculate the final complete multiplication result in multiple steps.
[0145] In an optional implementation manner of this embodiment, the data of each matrix in the compressed matrix set can be processed in chunks by combining the special data structure of the compressed matrix set. At the same time, the data of the second associated matrix can be processed in chunks by combining the selection array matching the second sparse matrix, so as to improve the data transfer and subsequent multiplication calculation efficiency.
[0146] Generally, a computing chip contains multiple sparse computing units. Inside each sparse computing unit, one or more threads can be started to perform corresponding sparse computing tasks. By using the cooperation of the sparse computing units in the computing chip, partial data in the first sparse matrix and the second associated matrix can be loaded in multiple steps from each hardware register for local calculation, and the local calculation results can be combined or accumulated to finally obtain the multiplication result of the first sparse matrix and the second associated matrix.
[0147] In the calculation process of the mixture of experts model, there are many one-to-one operators adjacent to matrix multiplication, but the calculation of two consecutive operators will introduce redundant computing core startup overhead and I / O (Input / Output) overhead. To eliminate the overhead of loading and storing intermediate results and multiple kernel startups, the calculation cores of each sparse operator can be improved by using the method of operator fusion.
[0148] Among various sparse operators, the first sparse operator is usually the operator that outputs the calculation result of the model. Exemplarily, Figure 8 as shown, the sparse operator 3 in the target mixture of experts model can be used as the first sparse operator. Optionally, for the first sparse operator, after the multiplication result of the first sparse matrix and the second associated matrix is gradually calculated by the sparse computing unit of the computing chip according to the data loaded in multiple steps from the hardware registers, the multiplication result of the first sparse matrix and the second associated matrix can be processed according to the processing logic of the broadcast multiplication function of the original mixture of experts model to obtain the first sparse operator with broadcast function.
[0149] Such as Figure 8As shown, the sparse operator 1 in the target mixture-of-experts model can be used as the second sparse operator. Optionally, for the second sparse operator, after the multiplication result of the first sparse matrix and the second correlation matrix is gradually calculated by the sparse computing unit of the computing chip according to the data loaded in batches from the hardware register, the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the activation function of the original mixture-of-experts model to obtain the second sparse operator. Accordingly, the second sparse operator can directly output the calculation result of the activation function.
[0150] As Figure 8 shown, the sparse operator 2 in the target mixture-of-experts model can be used as the third sparse operator. Optionally, for the third sparse operator, after the multiplication result of the first sparse matrix and the second correlation matrix is gradually calculated by the sparse computing unit of the computing chip according to the data loaded in batches from the hardware register, the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the dot product function of the original mixture-of-experts model to obtain the third sparse operator. Accordingly, the third sparse operator can directly output the calculation result of the dot product function.
[0151] For the first sparse operator, the second sparse operator, and the third sparse operator, the overall generation process of the sparse operator is generally the same. The only difference is that in the process of gradually transferring data from the global memory to the hardware register of the computing chip in the form of data blocks, the second sparse operator and the third sparse operator need to rearrange or transform the data in the original global memory during the transfer process. At the same time, because the operator functions are different, the types of specific functional functions fused when the first sparse operator, the second sparse operator, and the third sparse operator perform operator fusion are different.
[0152] S330. Construct a sparse operator library according to the first sparse operator, the second sparse operator, and the third sparse operator.
[0153] Among them, the sparse operator library can be used to store various generated sparse operators.
[0154] In the embodiments of the present invention, as Figure 8As shown, after generating various sparse operators suitable for the original mixture-of-experts model by adopting the data chunking method in combination with the sparse computing unit of the computing chip, the generated various sparse operators can be stored in the sparse operator library. It can be understood that the above process of generating sparse operators only schematically illustrates the specific generation processes of the first sparse operator, the second sparse operator, and the third sparse operator. For other types of sparse operators in the original mixture-of-experts model, similar methods as above can also be used for optimized generation. That is to say, the sparse operators are not limited to the first sparse operator, the second sparse operator, and the third sparse operator. Correspondingly, in addition to storing the first sparse operator, the second sparse operator, and the third sparse operator, the sparse operator library can also store other types of sparse operators, and the embodiments of the present invention do not limit the types of sparse operators.
[0155] S340. Screen the target sparse operator from the sparse operator library according to the type of the target mixture-of-experts model, and update the original operator of the target mixture-of-experts model according to the target sparse operator to obtain a structured sparse mixture-of-experts model.
[0156] Wherein, the target sparse operator can be a sparse operator used to update and replace the original operator in the original mixture-of-experts model. Optionally, the number of target sparse operators can be multiple.
[0157] Exemplarily, as Figure 8 shown, the original operators of the target mixture-of-experts model can include a first matrix multiplication operator, a second matrix multiplication operator, and a third matrix multiplication operator. Among them, the first matrix multiplication operator can be the matrix multiplication 2 operator in the target mixture-of-experts model before optimization, the second matrix multiplication operator can be the matrix multiplication 1 operator in the target mixture-of-experts model before optimization, and the third matrix multiplication operator can be the matrix multiplication 3 operator in the target mixture-of-experts model before optimization. For this type of target mixture-of-experts model, the first sparse operator, the second sparse operator, and the third sparse operator can be screened as the target sparse operators. Correspondingly, when updating the original operators of the target mixture-of-experts model according to the target sparse operators, the first matrix multiplication operator of the target mixture-of-experts model can be replaced and updated according to the first sparse operator, the second matrix multiplication operator of the target mixture-of-experts model can be replaced and updated according to the second sparse operator, and the third matrix multiplication operator of the target mixture-of-experts model can be replaced and updated according to the third sparse operator to obtain a structured sparse mixture-of-experts model.
[0158] Exemplarily, as Figure 8As shown, the original operators of some types of target mixture-of-experts models only include a first matrix multiplication operator and a second matrix multiplication operator. Among them, the first matrix multiplication operator can be the matrix multiplication 2 operator in the target mixture-of-experts model before optimization, and the second matrix multiplication operator can be the matrix multiplication 1 operator in the target mixture-of-experts model before optimization. For this type of target mixture-of-experts model, a first sparse operator and a second sparse operator can be selected as target sparse operators. Correspondingly, when updating the original operators of the target mixture-of-experts model according to the target sparse operators, the first matrix multiplication operator of the target mixture-of-experts model can be replaced and updated according to the first sparse operator, and the second matrix multiplication operator of the target mixture-of-experts model can be replaced and updated according to the second sparse operator, so as to obtain a structured sparse mixture-of-experts model.
[0159] Similarly, for other types of target mixture-of-experts models, sparse operator types with other functions can also be adaptively generated, and the original operators inside can be replaced and updated to obtain the corresponding structured sparse mixture-of-experts models.
[0160] The technical solution of the embodiments of the present invention, after compressing and storing the sparse matrix that needs to perform multiplication calculations in a specific compression format, transfers the sparse matrix in the above specific compression format and the associated matrix to the hardware registers of the computing chip level by level in the form of data blocks from the global memory, and through the sparse computing unit of the computing chip, according to the data loaded in batches in the hardware registers, gradually calculates the multiplication result of the sparse matrix and the associated matrix. This implementation method fully considers the sparse characteristics when performing matrix multiplication calculations based on the sparse matrix. By adopting data block and data transfer technologies adapted to this sparse characteristic, the hardware acceleration performance of the sparse computing unit in the computing chip can be fully utilized, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the sparse matrix multiplication operation process, and is particularly suitable for the model calculation scenario of the mixture-of-experts model.
[0161] Figure 9 It is a flowchart of another optimization method for the mixture-of-experts model provided by the embodiments of the present invention. This embodiment is optimized based on the above-mentioned embodiments. In this embodiment, the operation of "transferring the compressed matrix set and the second associated matrix to the hardware registers of the computing chip level by level in the form of data blocks" is specifically implemented.
[0162] Correspondingly, as Figure 9 shown, this method may include:
[0163] S410. Perform sparse conversion on various sparse data structures of the target mixture-of-experts model to obtain a structured sparse data structure, and optimize the data arrangement of the structured sparse data structure.
[0164] S420. Locate the set of compressed matrices and the second associated matrix that match the first sparse matrix in the global memory of the computing chip according to the sparse matrix multiplication requirement.
[0165] Among them, the sparse matrix multiplication requirement may include, but is not limited to, the first sparse matrix multiplication requirement, the second sparse matrix multiplication requirement, the third sparse matrix multiplication requirement, etc.
[0166] In an optional implementation manner of the embodiment of the present invention, the first sparse matrix may be the sparsified model weight matrix in the target mixture-of-experts model, the second sparse matrix may be the intermediate activation value sparse matrix input to the mixture model layer in the target mixture-of-experts model, and the dense matrix corresponding to the second sparse matrix may be the dense matrix of the intermediate activation value input to the mixture model layer in the target mixture-of-experts model.
[0167] In the computing scenario of the mixture-of-experts model, by performing the sparse matrix multiplication process based on the specific model weight matrix and the matching sparse matrix or dense matrix of the intermediate activation value, and combining the operator fusion strategy, various types of sparse operators can be generated to realize the optimization of the operators of the mixture-of-experts model, and further complete the optimization of the internal model structure of the mixture-of-experts model.
[0168] S430. Perform data chunking on the compressed data matrix in the global memory, and perform matching data chunking on the index matrix and the metadata matrix according to the data correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix.
[0169] In the embodiment of the present invention, since the compressed data matrix stores all non-zero data in the first sparse matrix, data chunking can be performed according to the data scale of the compressed data matrix. In the process of performing data chunking on the compressed data matrix, in order to further consider the efficiency of subsequent data transfer and multiplication calculation, it can be set that during the process of performing data chunking on the compressed data matrix in the global memory, V (the number of columns included in each compression unit) is an integer multiple of the chunking size in the column splitting direction.
[0170] For example, assume that in the global memory, when the compressed data matrix is sliced into multiple data chunks of size m b *k b it is required that V be an integer multiple of k b For example, k b is V / 2 or V / 4, etc.
[0171] As shown above, since there is a data or position correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix. After determining the compressed data matrix blocks sliced from the compressed data matrix, the index matrix blocks and metadata matrix blocks matching the compressed data matrix blocks can be uniquely determined. That is, according to the data correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix, the data blocks of the index matrix and the metadata matrix are matched.
[0172] S440. Locate the selection array matching the second correlation matrix in the computing chip, and according to the selection array, perform data blocking on the second correlation matrix in the global memory.
[0173] Among them, the selection array is used to describe the valid row positions in the second sparse matrix. The selection array of the second sparse matrix is the index of the non-zero data positions, and the selection array of the dense matrix corresponding to the second sparse matrix is the position mapping of the data in the dense matrix in the second sparse matrix.
[0174] As shown above, if the second correlation matrix is a second sparse matrix, there may be a large number of sparse rows in the second sparse matrix. If these sparse rows are also data-blocked and sent to the hardware computing unit for matrix multiplication calculation, it will bring redundant data loading and calculation, reducing the calculation efficiency.
[0175] Based on this, in the embodiments of the present invention, in the global memory, based on the valid row positions defined by the selection array, only the valid data rows in the second correlation matrix are data-blocked. Optionally, one or more valid data rows can be taken as a data block each time, or a set number of data in one valid data row can be taken as a data block each time, etc.
[0176] S450. Load each compressed data matrix block, each index matrix block, and each metadata matrix block from the global memory into the shared memory successively.
[0177] After completing the data blocking of the compressed matrix set in the global memory, the above-mentioned compressed data matrix blocks, index matrix blocks, and metadata matrix blocks can be loaded from the global memory into the shared memory successively.
[0178] On the basis of the above embodiments, in order to prevent the data in the shared memory from being restricted by the bank conflict that the same memory bank (also called bank) in the computing chip cannot be accessed by multiple threads simultaneously during the access (read / write) process, in the embodiments of the present invention, it is further considered to re-arrange the storage of data in the shared memory using a specific data rearrangement function based on the data offset.
[0179] Correspondingly, in an alternative implementation of the embodiment of the present invention, after successively transferring each compressed data matrix block, each index matrix block, and each metadata matrix block from the global memory to the shared memory, it may further include:
[0180] Using a preset data rearrangement function to rearrange each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory to avoid bank conflicts.
[0181] Specifically, according to the specific type and model of the computing chip, a matching data rearrangement function can be selected from the adapted function library to rearrange the data in the shared memory, and the embodiment of the present invention does not limit this.
[0182] S460. Successively transfer each second correlation matrix block from the global memory to the shared memory.
[0183] Similarly, after completing the data block division of the second correlation matrix in the global memory, each second correlation matrix block can be successively transferred from the global memory to the shared memory accordingly.
[0184] Meanwhile, to avoid the problem of bank conflicts, in an alternative implementation of the embodiment of the present invention, after successively transferring each second correlation matrix block from the global memory to the shared memory, it may further include:
[0185] Using a preset data rearrangement function to rearrange each second correlation matrix block in the shared memory to avoid bank conflicts.
[0186] Of course, it can be understood that in addition to using the data rearrangement function to rearrange each compressed data matrix block, each index matrix block, each metadata matrix block, and each second correlation matrix block in the shared memory, it is also possible to add data padding to each compressed data matrix block, each index matrix block, each metadata matrix block, and each second correlation matrix block in the shared memory to avoid bank conflicts.
[0187] S470. Perform secondary block division on each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory, and perform secondary block division on each second correlation matrix block in the shared memory.
[0188] In the embodiment of the present invention, in order to adapt to the storage limit of the hardware register, secondary block division can be performed on each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory.
[0189] Further, when performing secondary block division on each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory, the secondary block division of the compressed data matrix block can also be performed first. During the secondary block division of the compressed data matrix block, to further consider the efficiency of subsequent data transfer and multiplication calculations, it can be set that during the secondary block division of the compressed data matrix block in the shared memory, V is also an integer multiple of the division size in the column splitting direction. For example, secondary splitting can be selected in the column splitting direction or no secondary splitting can be selected (only splitting in the row direction).
[0190] For example, assume that in the global memory, when the compressed data matrix block is secondarily split into multiple new data blocks of size m b1 *k b1 it is required that V is an integer multiple of k b1 For example, k b1 is V / 4 or V / 8, etc.
[0191] Similarly, since there is a data or position correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix, after the secondary block division of the compressed data matrix block is completed, the secondary block division of the index matrix block and the metadata matrix block can be correspondingly performed.
[0192] S480. According to the register mapping relationship matching the sparse computing unit, at least one of the secondary block divisions of each compressed data matrix, the secondary block divisions of each index matrix, and the secondary block divisions of each metadata matrix is sequentially transferred from the shared memory to the hardware registers.
[0193] For a more intuitive understanding, a block diagram of the implementation of a data block division and a multi-level transfer process applicable to the embodiments of the present invention is shown in Figure 10 . In a specific example, taking the second sparse matrix as an example, as shown in Figure 10 , it describes that after the compressed data matrix A in the first sparse matrix and the second sparse matrix B are respectively divided into data blocks in the global memory and the shared memory, and are successively transferred in the order of global memory -> shared memory -> register, after the calculation is completed by the sparse computing unit and gradually written back, the corresponding multiplication result matrix C is obtained in the global memory.
[0194] In this example, an example of data block division and data transfer for matrix multiplication of any size C(M×N)=A(M×K)×B(K×N) is given. In the embodiments of the present invention, in order to implement matrix multiplication calculation, the second sparse matrix B is row-column interchanged, that is, an array is selected to describe the positions where the valid data columns are located in the second sparse matrix B.
[0195] Specifically, in the global memory, since only the positions where the valid data columns (rows) are located are included in the selection array, the total amount of data in the selection array is len d and thus the number of valid values in the n dimension of the second sparse matrix B is len d The compressed data matrix A is evenly divided into several parts in the m dimension with m b as the size. The second sparse matrix B is divided in the n dimension, and the number of valid column vectors in each block of data is the same as the data block divided with the size of n b in the selection array. Considering the storage size limitation of the computing device, the compressed data matrix A and the second sparse matrix B will be further evenly divided along the k direction with k b as the size. The multiplication results of multiple sub-blocks in the k direction are accumulated to obtain the final multiplication result matrix C block result. Based on this data partitioning strategy, the size of the multiplication result matrix C calculated for each block of data is m b ×n b , and the result calculated by block is a part of the final result of the multiplication result matrix C, corresponding to the same row offset of the data block of the compressed data matrix A in the compressed data matrix A and the position where the data block of the second sparse matrix B has the same column offset in the selection array. The parallelism of each sparse operator is improved by having each thread group be responsible for one block.
[0196] Similarly, in the shared memory, it also involves the process of re-blocking each data block transferred from the global memory, moving it to the register for storage, and then performing matrix multiplication calculations by the adapted sparse computing unit, which will not be elaborated here. In this embodiment, the matrix specifications of the compressed data matrix A processed by the sparse computing unit at one time are mi * ki, and the matrix specifications of the second sparse matrix B are ki * ni. For more convenient subsequent calculations, the aforementioned V is an integer multiple of ki.
[0197] In an alternative implementation manner of the embodiments of the present invention, the re-blocking of each compressed data matrix, the re-blocking of each index matrix, and the re-blocking of each metadata matrix can all be moved to the register for implementation of calculations. Or, considering that register resources are very precious, only the re-blocking of each compressed data matrix and the re-blocking of each metadata matrix can be moved to the register for implementation of calculations, while the index matrix is retained in the shared memory and only obtained from the shared memory for use when needed.
[0198] It should be emphasized that different from the calculation logic of general hardware computing units, when using a sparse computing unit to accelerate multiplication calculations, the data required for calculation needs to be placed in the adapted instruction register according to the requirements of the sparse instruction set. Therefore, in the embodiments of the present invention, the data in the shared memory needs to be moved to the matching hardware register according to the register mapping relationship matching the sparse computing unit.
[0199] Optionally, the register mapping relationship can be read from the specification file configured at the factory of the computing chip, and the register hardware file describes the mapping relationship between the computing data at different positions and the hardware registers.
[0200] Specifically, in Figure 11 a schematic diagram of a register mapping relationship applicable to the embodiments of the present invention is shown, and this register mapping relationship is adapted to the quadratic sub-blocking of the compressed data matrix. In a specific example, as Figure 11 shown, T0{a0, a1} at the upper left corner position in this register mapping relationship represents: the data in the 0th and 1st columns of the 0th row in the quadratic sub-blocking of the compressed data matrix are allocated to the register a0 and register a1 that match the thread T0. Similarly, T 0…3 {a4, a5} at the last position in the 0th row represents: the data in the 8th - 15th columns of the 0th row in the quadratic sub-blocking of the compressed data matrix are respectively allocated to the register a4 and register a5 that match the threads T0, T1, T2, and T3. Among them, different threads T are pre-allocated to different sparse computing units for use.
[0201] Similarly, for the quadratic sub-blocking of the metadata matrix and the quadratic sub-blocking of the index matrix, the sparse computing unit also has a register mapping relationship that is adapted. Based on the above register mapping relationship, each item of computing data can be successively moved from the shared memory to the hardware registers.
[0202] In an optional implementation manner of the embodiments of the present invention, according to the register mapping relationship that matches the sparse computing unit, a preset hardware instruction can be called to successively move each quadratic sub-block of the compressed data matrix and each quadratic sub-block of the metadata matrix, or each quadratic sub-block of the compressed data matrix, each quadratic sub-block of the index matrix, and each quadratic sub-block of the metadata matrix from the shared memory to the hardware registers to achieve hardware acceleration of data loading.
[0203] Among them, this hardware instruction is associated with the specification of the computing chip and can be queried and obtained from the instruction library of this computing chip. For example, it can be a 1dmatrix instruction, etc.
[0204] S490. According to the register mapping relationship that matches the sparse computing unit, successively load each quadratic sub-block of the second associated matrix from the shared memory to the hardware registers.
[0205] As shown before, by querying and obtaining the register mapping relationship set by the sparse computing unit for the second sparse matrix quadratic sub-block, each quadratic sub-block of the second sparse matrix can be successively loaded from the shared memory to the hardware registers.
[0206] Similarly, in an alternative embodiment of the present invention, according to the register mapping relationship matching the sparse computing unit, a preset hardware instruction can be called to secondarily partition each second sparse matrix and sequentially load it from the shared memory into the hardware register to achieve hardware acceleration of data loading.
[0207] It should be noted that in addition to the first sparse operator, when the second sparse operator and the third sparse operator are sequentially transferred from the global memory to the hardware register of the computing chip in the form of data blocks, during the transfer process, it is necessary to rearrange or transform the data in the original global memory.
[0208] S4100. Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second correlation matrix according to the data loaded in batches in the hardware register.
[0209] S4110. Process the multiplication result of the first sparse matrix and the second correlation matrix according to the processing logic of the corresponding functional function of the original mixture-of-experts model to obtain a sparse operator.
[0210] S4120. Update the original operator of the target mixture-of-experts model according to the sparse operator to obtain a structured sparse mixture-of-experts model.
[0211] Optionally, the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the broadcast multiplication function of the original mixture-of-experts model to obtain the first sparse operator; the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the activation function of the original mixture-of-experts model to obtain the second sparse operator; the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the dot product function of the original mixture-of-experts model to obtain the third sparse operator.
[0212] In an alternative embodiment of the present invention, gradually calculating the multiplication result of the first sparse matrix and the second sparse matrix through the sparse computing unit of the computing chip according to the data loaded in batches in the hardware register may include: obtaining, through the sparse computing unit, the current compressed data matrix secondary block, the current metadata matrix secondary block, and the current second correlation matrix secondary block currently loaded in the hardware register; generating, through the sparse computing unit, a primary sparse block matrix according to the current metadata matrix secondary block and the current compressed data matrix secondary block; and performing a multiplication calculation through the sparse computing unit according to the primary sparse block matrix and the current second correlation matrix secondary block.
[0213] As described above, when compressing and storing the first sparse matrix, the compressed data matrix only stores the non-zero data in each valid row. However, in the actual first sparse matrix, each valid row contains at least one structured sparse computing storage unit in a set sparse pattern. To ensure the accuracy of multiplication calculations, it is necessary to first decompose and restore the current compressed data matrix into the form of the aforementioned structured sparse computing storage units. Since the metadata matrix stores the positions of non-zero data in the corresponding structured sparse computing storage units, furthermore, based on the secondary block decomposition of the current metadata matrix and the secondary block decomposition of the current compressed data matrix, a first-level sparse block matrix can be generated.
[0214] That is to say, the first-level sparse block matrix can be understood as the matrix after restoring the structured sparse computing storage units in the secondary block decomposition of the current compressed data matrix. In a specific example, if the secondary block decomposition of the current compressed data matrix is {A, B}, and the corresponding secondary block decomposition of the current metadata matrix is {0, 2}, then a first-level sparse block matrix in the form of {A, 0, B, 0} can be restored.
[0215] Among them, through the sparse computing unit, according to the first-level sparse block matrix and the current secondary block decomposition of the second correlation matrix, perform multiplication calculations, which can specifically include: through the sparse computing unit, store the intermediate calculation amounts generated during the multiplication calculation process into the intermediate registers of the computing chip; and, through the sparse computing unit, perform the matching multiplication calculation by obtaining the intermediate calculation amounts from the intermediate registers.
[0216] In the prior art, when using a general computing unit to perform multiplication calculations, the results of matrix multiplication calculations can be directly accumulated, but in essence, it increases the calculation of many redundant data. In contrast, when the sparse computing units of the embodiments of the present invention perform sparse multiplication calculations, when the multiplication calculations iterate along the K direction, since the sparse data comes from different rows, to ensure correctness, when different sub-blocks move along the K direction, the output results need to be mapped to different rows. Traditionally, passing the output results to specific registers according to indexes may cause the output C matrix to be reloaded into the local memory, which has a significant impact on the performance of the operator.
[0217] To avoid the above problems, the embodiments of the present invention reduce the frequent memory transfer between the global memory and the registers by introducing an additional intermediate register C IR in this way. Among them, in Figure 12 shows a schematic structural diagram of a sparse computing unit cooperating with an intermediate register for calculation applicable to the embodiments of the present invention. As Figure 12 shown, for the secondary block decomposition of the current compressed data matrix ( Figure 12 A in), the secondary block decomposition of the current metadata matrix ( Figure 12the metadata therein) and the secondary block division of the previous second correlation matrix ( Figure 12 When performing sparse matrix multiplication calculation on B in IR for multiple result data, during the calculation process, load the required intermediate results into C
[0218] By adding a simple hardware improvement of a new intermediate register in the computing chip, during the process of performing sparse matrix multiplication calculation in the sparse computing unit, the computing performance can be greatly improved and the computing efficiency can be enhanced.
[0219] Based on the above embodiments, if the second correlation matrix includes a second sparse matrix, after performing multiplication calculation by the sparse computing unit according to the primary sparse block matrix and the current secondary block division of the second sparse matrix, it may further include: after determining that the complete multiplication calculation result is stored in the intermediate register, converting the complete multiplication calculation result into a secondary sparse block matrix according to the current secondary block division of the index matrix that matches the current compressed data matrix secondary block; writing back the secondary sparse block matrix to the global memory of the computing chip level by level.
[0220] If the second correlation matrix includes a dense matrix corresponding to the second sparse matrix, after performing multiplication calculation by the sparse computing unit according to the primary sparse block matrix and the current secondary block division of the second correlation matrix, it may further include: after determining that the complete multiplication calculation result is stored in the intermediate register, writing back the multiplication calculation result to the global memory of the computing chip level by level in the original sparse format.
[0221] Furthermore, since the sparse rows in the first sparse matrix are not considered when obtaining the complete multiplication calculation result, therefore, if the right operand uses the second sparse matrix, after obtaining the complete multiplication calculation result, in combination with the current secondary block division of the index matrix that matches the current compressed data matrix secondary block, add the matching sparse rows to the complete multiplication calculation result to obtain a secondary sparse block matrix as the true block calculation result. As mentioned above, the current secondary block division of the index matrix can also be transferred to the hardware register or only stored in the shared memory, and this embodiment does not limit this. If the right operand uses the dense matrix corresponding to the second sparse matrix, after obtaining the complete multiplication calculation result, the multiplication calculation result can be directly written back to the global memory of the computing chip level by level in the original sparse format without the need to perform sparse processing on the multiplication calculation result. As mentioned above, the multiplication calculation result can be transferred to the hardware register in the original sparse format or only stored in the shared memory, and this embodiment does not limit this.
[0222] In summary, in combination with the model architecture, matrix multiplication operations with sparse attributes are selected from the calculation flow of the mixture-of-experts model. Sparse comes from two aspects. On the one hand, it comes from the sparsity of the weights. The redundant attributes in its parameters enable the calculation process to skip the calculation of less important parameters. On the other hand, sparsity can also come from the mixture-of-experts model. The input values are selected through a gating circuit to match the expert parameters for calculation. From the perspective of the expert, the input values have sparse characteristics, and only part of the input participates in the calculation of the current expert. The selection of such sparse operators effectively improves the execution speed of the mixture-of-experts model without affecting the model accuracy. Based on the strategy of operator fusion in sparse data structure and data stream execution, referring to the compilation optimization scheme of sparse operators, specific sparse operators in the sparse operator library are selected and injected into the artificial intelligence framework and compiler to replace the dense matrix multiplication operation of the original mixture-of-experts model. For example, the two matrix multiplication operations that receive the original input in the expert architecture of the mixture-of-experts model will be injected with operators in the double-sparse mode, while the subsequent matrix multiplication that receives the intermediate activation values will be injected with unidirectional sparse (sparse × dense) operators.
[0223] The technical solution of the embodiment of the present invention fully considers the sparse characteristics when performing matrix multiplication calculations based on sparse matrices. By adopting data partitioning and data transfer techniques adapted to the sparse characteristics, operator optimization of the mixture-of-experts model is achieved, and the hardware acceleration performance of the sparse computing units in the computing chip is fully utilized in the model calculation scenario of the mixture-of-experts model, thereby greatly optimizing the overhead of computing, bandwidth, and storage resources in the sparse matrix multiplication operation process.
[0224] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards of the relevant regions.
[0225] It should be noted that any permutation and combination of the technical features in the above embodiments also fall within the protection scope of the present invention.
[0226] Figure 13 is a schematic diagram of an optimization device for a mixture-of-experts model provided by an embodiment of the present invention. As Figure 13 shown, the device includes: a data structure sparse conversion module 510, a structured sparse data structure optimization module 520, a sparse operator optimization module 530, and a sparse operator update module 540, where:
[0227] The data structure sparse conversion module 510 is used to perform sparse conversion on various sparse data structures of the target mixture-of-experts model to obtain a structured sparse data structure;
[0228] A structured sparse data structure optimization module 520 is configured to optimize the data arrangement of the structured sparse data structure;
[0229] A sparse operator optimization module 530 is configured to optimize the operators of the target mixture-of-experts model in combination with the sparse computing units of the computing chip to obtain sparse operators;
[0230] A sparse operator update module 540 is configured to update the original operators of the target mixture-of-experts model according to the sparse operators to obtain a structured sparse mixture-of-experts model.
[0231] In the embodiment of the present invention, by performing sparse conversion on various sparse data structures of the target mixture-of-experts model, a structured sparse data structure is obtained, and the data arrangement of the structured sparse data structure is optimized. Further, the operators of the target mixture-of-experts model are optimized in combination with the sparse computing units of the computing chip to obtain sparse operators, and then the original operators of the target mixture-of-experts model are updated according to the sparse operators to obtain a structured sparse mixture-of-experts model. The above technical solution can solve the problem that it is difficult to effectively improve the performance when the existing mixture-of-experts model performs performance optimization and acceleration based on the unstructured sparse computing technology, and can realize the structured sparse optimization of the mixture-of-experts model to give full play to the acceleration performance of the sparse computing hardware of the mixture-of-experts model when calculating the sparse matrix multiplication, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the sparse matrix multiplication operation process.
[0232] Optionally, the data structure sparse conversion module 510 is further configured to: load the target mixture-of-experts model into the computing chip and obtain various sparse model weight matrices in the target mixture-of-experts model; wherein, the sparse model weight matrix includes a plurality of compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; generate a compression matrix set corresponding to each model weight matrix respectively; wherein, the compression matrix set includes a compressed data matrix for storing the non-0 data in the model weight matrix, an index matrix for storing the positions of the valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of the non-0 data in the corresponding structured sparse computing storage unit; store the various sparse model weight matrices in the target mixture-of-experts model as the matching compression matrix sets respectively to obtain a structured sparse data structure of the model weight data.
[0233] Optionally, the sparsified model weight matrix specifically includes multiple compression units of size M*V; each of the compression units specifically includes d valid rows, where 1≤d<M; and each of the structured sparse computing and storage units is in an N:L sparse mode, where L is the total number of data included in the structured sparse computing and storage unit, and N is the number of non-zero data included in the structured sparse computing and storage unit.
[0234] Optionally, the data structure sparse conversion module 510 is further configured to: obtain a current model weight matrix of size m*k being processed, where m is an integer multiple of M and k is an integer multiple of V; construct a first matrix of size (m / M*d)*(k / (L / N)), and fill each non-zero element of the current model weight matrix into the first matrix in a row-by-row traversal manner to obtain a compressed data matrix corresponding to the current model weight matrix; construct a second matrix of size (m / M*d)*(k / V), and perform a filling process on the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix; construct a third matrix of size (m / M*d)*(k / (L / N)), and perform a filling process on the third matrix according to the positions of each of the structured sparse computing and storage units in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing and storage unit to obtain a metadata matrix corresponding to the current model weight matrix.
[0235] Optionally, the data structure sparse conversion module 510 is further configured to: sequentially traverse a current compression unit in the current model weight matrix, and locate d longitudinal matrix positions in the second matrix that match the current compression unit according to the position of the current compression unit in the current model weight matrix; identify the row positions where each of the valid rows is located in the current compression unit, and fill the identified row positions into the d longitudinal matrix positions; return to perform the operation of sequentially traversing a current compression unit in the current model weight matrix until the processing of all compression units in the current model weight matrix is completed, so as to obtain the index matrix corresponding to the current model weight matrix.
[0236] Optionally, the data structure sparse conversion module 510 is further configured to: sequentially traverse a current structured sparse computing storage unit in the current model weight matrix, and locate N horizontal matrix positions matching the current structured sparse computing storage unit in the third matrix according to the position of the current structured sparse computing storage unit in the current model weight matrix; fill the column positions where each non-0 data is located in the current structured sparse computing storage unit into the N horizontal matrix positions correspondingly; return to perform the operation of sequentially traversing a current structured sparse computing storage unit in the current model weight matrix until the processing of all the structured sparse computing storage units in the current model weight matrix is completed, so as to obtain the metadata matrix corresponding to the current model weight matrix.
[0237] Optionally, the above device further includes a sparse computing module, configured to: call each of the sparse operators on the computing chip, and perform matching sparse computing based on the target mixture-of-experts model after storage optimization.
[0238] Optionally, the above device further includes a matrix set rearrangement storage module, configured to: obtain a sparse computing unit requirement file matching the computing chip; wherein, the sparse computing unit requirement file defines the data access positions of each thread executed in the sparse operator to each model weight matrix in the target mixture-of-experts model during sparse computing; rearrange the data arrangement modes of the compressed data matrix, the index matrix, and the metadata matrix in each of the compressed matrix sets according to the sparse computing unit requirement file; store the rearranged compressed matrix sets in the set memory of the computing chip, so as to realize continuous storage of data with continuous access characteristics.
[0239] Optionally, the sparse computing module is further configured to: identify each mixture model layer included in the target mixture-of-experts model, and obtain an expert allocation mode corresponding to each mixture model layer; wherein, the expert allocation mode is used to indicate the row position where the valid rows of the intermediate activation values allocated to its own mixture model layer are located during calculation; generate a selection array corresponding to each mixture model layer according to the expert allocation mode; and associatively store the selection arrays of each mixture model layer with the compressed matrix sets of each model weight matrix, so as to combine the selection array with the sparse matrix or dense matrix of the intermediate activation values to be processed during calculation and perform sparse computing with the matching compressed matrix set.
[0240] Optionally, the structured sparse data structure optimization module 520 is further configured to: before running the target mixture-of-experts model, store the sparsely transformed model weight data of the target mixture-of-experts model in a transposed manner; when the sparse operator of the target mixture-of-experts model performs data operations in shared memory, perform a data transpose operation on the intermediate activation values.
[0241] Optionally, the sparse operator includes a first sparse operator, and the sparse operator optimization module 530 is further configured to: according to the first sparse matrix multiplication requirement, locate a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; transfer the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in the form of data blocks level by level; through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the broadcast multiplication function of the target mixture-of-experts model to obtain the first sparse operator.
[0242] Optionally, the sparse operator includes a second sparse operator, and the sparse operator optimization module 530 is further configured to: according to the second sparse matrix multiplication requirement, locate a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; transfer the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in the form of data blocks level by level, and during the transfer process, rearrange or deform the data in the original global memory; through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the activation function of the target mixture-of-experts model to obtain the second sparse operator.
[0243] Optionally, the sparse operator includes a third sparse operator, and the sparse operator optimization module 530 is further configured to: locate, according to the third sparse matrix multiplication requirement, a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; transfer the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in the form of data chunks, and rearrange or transform the data in the original global memory during the transfer process; through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the dot product function of the target mixture-of-experts model to obtain the third sparse operator.
[0244] Optionally, the first sparse matrix is the sparsified model weight matrix in the target mixture-of-experts model, the second sparse matrix is the intermediate activation value sparse matrix input to the mixture model layer in the target mixture-of-experts model, and the dense matrix corresponding to the second sparse matrix is the dense matrix of the intermediate activation value input to the mixture model layer in the target mixture-of-experts model.
[0245] Optionally, the sparse operator optimization module 530 is further configured to: divide the compressed data matrix in the global memory into data chunks, and divide the index matrix and the metadata matrix into matching data chunks according to the data correspondence relationship between the compressed data matrix, the index matrix, and the metadata matrix; sequentially load each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk from the global memory to the shared memory; perform secondary chunking on each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk in the shared memory; according to the register mapping relationship matching the sparse computing unit, sequentially transfer at least one of each compressed data matrix secondary chunk, each index matrix secondary chunk, and each metadata matrix secondary chunk from the shared memory to the hardware registers.
[0246] Optionally, the sparse operator optimization module 530 is further configured to: locate a selection array that matches the second association matrix in the computing chip, where the selection array is used to describe valid row positions in the second sparse matrix; the selection array of the second sparse matrix is the index of the non-zero data positions therein, and the selection array of the dense matrix corresponding to the second sparse matrix is the position mapping of the data in the dense matrix in the second sparse matrix; according to the selection array, perform data chunking on the second association matrix in the global memory; for each chunk of the second association matrix, sequentially transfer it from the global memory to the shared memory; perform secondary chunking on each chunk of the second association matrix in the shared memory; according to the register mapping relationship that matches the sparse computing unit, for each secondary chunk of the second association matrix, sequentially load it from the shared memory to the hardware register.
[0247] Optionally, the sparse operator optimization module 530 is further configured to: use a preset data rearrangement function to rearrange each chunk of the compressed data matrix, each chunk of the index matrix, and each chunk of the metadata matrix in the shared memory to avoid bank conflicts; use the preset data rearrangement function to rearrange each chunk of the second association matrix in the shared memory to avoid bank conflicts.
[0248] Optionally, the sparse operator optimization module 530 is further configured to: according to the register mapping relationship that matches the sparse computing unit, call a preset hardware instruction to perform secondary chunking on each chunk of the compressed data matrix and each chunk of the metadata matrix, or perform secondary chunking on each chunk of the compressed data matrix, each chunk of the index matrix, and each chunk of the metadata matrix, and sequentially transfer them from the shared memory to the hardware register to achieve hardware acceleration of data loading; according to the register mapping relationship that matches the sparse computing unit, call a preset hardware instruction to perform secondary chunking on each chunk of the second association matrix and sequentially load it from the shared memory to the hardware register to achieve hardware acceleration of data loading.
[0249] Optionally, the sparse operator optimization module 530 is further configured to: through the sparse computing unit, obtain the current secondary chunk of the compressed data matrix, the current secondary chunk of the metadata matrix, and the current secondary chunk of the second association matrix currently loaded in the hardware register; through the sparse computing unit, generate a first sparse chunk matrix according to the current secondary chunk of the metadata matrix and the current secondary chunk of the compressed data matrix; through the sparse computing unit, perform multiplication calculation according to the first sparse chunk matrix and the current secondary chunk of the second association matrix.
[0250] Optionally, the sparse operator optimization module 530 is further configured to: store the intermediate computation amount generated during the multiplication computation into the intermediate register of the computing chip through the sparse computing unit; and, execute the matching multiplication computation by obtaining the intermediate computation amount from the intermediate register through the sparse computing unit.
[0251] Optionally, if the second correlation matrix includes a second sparse matrix, the sparse operator optimization module 530 is further configured to: after determining that the complete multiplication computation result is stored in the intermediate register, convert the complete multiplication computation result into a quadratic sparse block matrix according to the current index matrix quadratic block that matches the quadratic block division of the current compressed data matrix; write back the quadratic sparse block matrix to the global memory of the computing chip level by level; if the second correlation matrix includes the dense matrix corresponding to the second sparse matrix, the sparse operator optimization module 530 is further configured to: after determining that the complete multiplication computation result is stored in the intermediate register, write back the multiplication computation result to the global memory of the computing chip level by level in the original sparse format.
[0252] Optionally, during each data block division, V is an integer multiple of the division size in the column division direction; and, in each sparse operator, the matrix specification of the first sparse matrix processed once is mi*ki, where V is an integer multiple of ki.
[0253] The above-mentioned hybrid expert model optimization device can execute the hybrid expert model optimization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the hybrid expert model optimization method provided in any embodiment of the present invention.
[0254] Since the above-mentioned hybrid expert model optimization device is a device that can execute the hybrid expert model optimization method in the embodiments of the present invention, based on the hybrid expert model optimization method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the hybrid expert model optimization device in this embodiment. Therefore, the details of how the hybrid expert model optimization device implements the hybrid expert model optimization method in the embodiments of the present invention will not be described in detail here. As long as the device adopted by those skilled in the art to implement the hybrid expert model optimization method in the embodiments of the present invention belongs to the scope to be protected by this application.
[0255] Figure 14The structural schematic diagram of an electronic device 10 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0256] As Figure 14 shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0257] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0258] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the optimization method of the mixture-of-experts model.
[0259] Optionally, the optimization method of the mixture-of-experts model may include: performing sparse conversion on various sparsified data structures of the target mixture-of-experts model to obtain a structured sparse data structure; optimizing the data layout of the structured sparse data structure; optimizing the operators of the target mixture-of-experts model in combination with the sparse computing units of the computing chip to obtain sparse operators; and updating the original operators of the target mixture-of-experts model according to the sparse operators to obtain a structured sparse mixture-of-experts model.
[0260] In some embodiments, the optimization method of the mixture-of-experts model may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the optimization method of the mixture-of-experts model described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the optimization method of the mixture-of-experts model in any other suitable manner (e.g., by means of firmware).
[0261] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, the programmable processor may be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0262] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to the processors of general-purpose computers, special-purpose computers, or other programmable data processing devices such that the computer programs, when executed by the processors, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0263] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0264] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0265] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0266] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0267] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.
[0268] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. An optimization method for a mixture of experts model, characterized in that include: Performing sparse conversion on the sparse data structures of each item of the target hybrid expert model to obtain a structured sparse data structure; wherein the structured sparse data structure includes a structured sparse data structure of model weight data, and the structured sparse data structure of the model weight data stores the sparse model weight matrices of each item in the target hybrid expert model through a compressed matrix set; Optimizing data arrangement of the structured sparse data structure; Optimizing the operator of the target hybrid expert model in combination with a sparse computing unit of a computing chip to obtain a sparse operator; Updating the original operator of the target hybrid expert model according to the sparse operator to obtain a structured sparse hybrid expert model; Obtaining a sparse computing unit requirement file matching the computing chip; The sparse computing unit requirement file defines the data access location of each thread executed in the sparse operator to each model weight matrix in the target hybrid expert model when each sparse computing unit on the computing chip performs sparse computing; According to the sparse computing unit requirement file, re-arrange the data arrangement of the compressed data matrix, the index matrix and the metadata matrix in each of the compressed matrix sets; The reorganized compression matrix sets are stored in a setting memory of a computing chip to achieve continuous storage of data with continuous access characteristics.
2. The method according to claim 1, wherein The sparse conversion of each sparse data structure of the target hybrid expert model is performed to obtain a structured sparse data structure, including: Loading the target hybrid expert model into a computing chip, and obtaining a sparse model weight matrix of each item in the target hybrid expert model; The sparse model weight matrix includes a plurality of compression units of set sizes, each of the compression units includes at least one valid row and at least one sparse row, each of the valid rows includes at least one structured sparse computing storage unit; each of the structured sparse computing storage units is in the same sparse mode; Generate a set of compression matrices corresponding to each of the model weight matrices; The compression matrix set includes a compression data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the position of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the position of non-zero data in the corresponding structured sparse computing storage unit; The sparse model weight matrices of each item in the target hybrid expert model are stored as matching compressed matrix sets respectively, so as to obtain a structured sparse data structure of the model weight data.
3. The method according to claim 2, wherein The sparse model weight matrix specifically includes a plurality of compression units of size M*V; Specifically, each compression unit includes d valid rows, wherein 1≤d<M; and Each of the structured sparse computing storage units is in a sparse mode of N:L, where L is the total amount of data contained in the structured sparse computing storage unit, and N is the amount of non-zero data contained in the structured sparse computing storage unit.
4. The method according to claim 3, wherein The generating a set of compressed matrices corresponding to each of the model weight matrices comprises: Obtain the current model weight matrix of size m*k being processed, where m is an integer multiple of M and k is an integer multiple of V; Construct a first matrix of size (m / M*d)*(k / (L / N)), and fill each non-zero element of the current model weight matrix into the first matrix row by row to obtain a compressed data matrix corresponding to the current model weight matrix; Construct a second matrix of size (m / M*d)*(k / V), and perform a filling process on the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix; Construct a third matrix of size (m / M*d)*(k / (L / N)), and perform a filling process on the third matrix according to the positions of each structured sparse computing and storage unit in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing and storage unit to obtain a metadata matrix corresponding to the current model weight matrix.
5. The method according to claim 4, wherein The performing a filling process on the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix includes: Traverse a current compression unit in the current model weight matrix in sequence, and locate d longitudinal matrix positions in the second matrix that match the current compression unit according to the position of the current compression unit in the current model weight matrix; Identify the row positions of each valid row in the current compression unit, and fill the identified row positions into the d longitudinal matrix positions correspondingly; Return to perform the operation of traversing a current compression unit in the current model weight matrix in sequence until the processing of all compression units in the current model weight matrix is completed to obtain the index matrix corresponding to the current model weight matrix.
6. The method according to claim 4, characterized in that The performing a filling process on the third matrix according to the positions of each structured sparse computing and storage unit in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing and storage unit to obtain a metadata matrix corresponding to the current model weight matrix includes: Traverse a current structured sparse computing and storage unit in the current model weight matrix in sequence, and locate N transverse matrix positions in the third matrix that match the current structured sparse computing and storage unit according to the position of the current structured sparse computing and storage unit in the current model weight matrix; Fill the column positions where each non-zero data is located in the current structured sparse computing and storage unit into the N transverse matrix positions correspondingly; Return to perform the operation of traversing a current structured sparse computing and storage unit in the current model weight matrix in sequence until the processing of all structured sparse computing and storage units in the current model weight matrix is completed to obtain the metadata matrix corresponding to the current model weight matrix.
7. The method according to any one of claims 2-6, characterized in that, After storing the model weight matrices of various sparsifications in the target mixture-of-experts model as a set of matching compressed matrices respectively, it further includes: Invoking each of the sparse operators on the computing chip to perform matching sparse computations based on the target mixture-of-experts model after storage optimization.
8. The method according to claim 1, wherein Before invoking each of the sparse operators on the computing chip to perform matching sparse computations based on the target mixture-of-experts model after storage optimization, it further includes: Identifying each mixture model layer included in the target mixture-of-experts model and obtaining an expert assignment pattern corresponding to each of the mixture model layers; Wherein, the expert assignment pattern is used to indicate the row position where the valid rows of the intermediate activation values assigned to its own mixture model layer are located during the implementation of the computation; Generating a selection array corresponding to each of the mixture model layers according to the expert assignment pattern; Associatively storing the selection arrays of each of the mixture model layers with the set of compressed matrices of each of the model weight matrices, so that during the implementation of the computation, the selection arrays are combined with the sparse matrix or dense matrix of the intermediate activation values to be processed and sparse computations are performed with the set of matching compressed matrices.
9. The method according to claim 1, wherein The data layout optimization for the structured sparse data structure includes: Before the target mixture-of-experts model runs, storing the model weight data after sparse transformation of the target mixture-of-experts model in a transposed manner; When the sparse operator of the target mixture-of-experts model performs data operations in the shared memory, performing a data transpose operation on the intermediate activation values.
10. The method according to claim 3, characterized in that, The sparse operator includes a first sparse operator. Optimizing the operator of the target mixture-of-experts model by combining the sparse computing unit of the computing chip to obtain a sparse operator, including: Locating a set of compressed matrices matching the first sparse matrix and a second associated matrix in the global memory of the computing chip according to the first sparse matrix multiplication requirement; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; Carrying the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data-block form step by step; Through the sparse computing unit of the computing chip, gradually calculating the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; Processing the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the broadcast multiplication function of the target mixture-of-experts model to obtain the first sparse operator.
11. The method according to claim 3, characterized in that, The sparse operator includes a second sparse operator. Optimizing the operator of the target mixture-of-experts model by combining the sparse computing unit of the computing chip to obtain a sparse operator, including: Locating a set of compressed matrices matching the first sparse matrix and a second associated matrix in the global memory of the computing chip according to the second sparse matrix multiplication requirement; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; Transfer the set of compressed matrices and the second correlation matrix to the hardware registers of the computing chip level by level in the form of data chunks from the global memory, and rearrange or transform the data in the original global memory during the transfer process; Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second correlation matrix according to the data loaded in batches in the hardware registers; Process the multiplication result of the first sparse matrix and the second correlation matrix according to the processing logic of the activation function of the target mixture-of-experts model to obtain the second sparse operator.
12. The method according to claim 3, wherein The sparse operator includes a third sparse operator. Optimizing the operator of the target mixture-of-experts model by combining the sparse computing unit of the computing chip to obtain a sparse operator includes: Locate a set of compressed matrices and a second correlation matrix that match the first sparse matrix in the global memory of the computing chip according to the third sparse matrix multiplication requirement; wherein, the second correlation matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; Transfer the set of compressed matrices and the second correlation matrix to the hardware registers of the computing chip level by level in the form of data chunks from the global memory, and rearrange or transform the data in the original global memory during the transfer process; Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second correlation matrix according to the data loaded in batches in the hardware registers; Process the multiplication result of the first sparse matrix and the second correlation matrix according to the processing logic of the dot product function of the target mixture-of-experts model to obtain the third sparse operator.
13. The method according to any one of claims 10 - 12, characterized in that, The first sparse matrix is the sparsified model weight matrix in the target mixture-of-experts model, the second sparse matrix is the intermediate activation value sparse matrix input to the mixture model layer in the target mixture-of-experts model, and the dense matrix corresponding to the second sparse matrix is the dense matrix of the intermediate activation value input to the mixture model layer in the target mixture-of-experts model.
14. The method according to claim 13, wherein Transferring the set of compressed matrices to the hardware registers of the computing chip level by level in the form of data chunks includes: Perform data chunking on the compressed data matrices in the global memory, and perform matching data chunking on the index matrices and the metadata matrices according to the data correspondence between the compressed data matrices, the index matrices, and the metadata matrices; Successively load each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk from the global memory to the shared memory; Perform secondary chunking on each of the compressed data matrix chunks, each of the index matrix chunks, and each of the metadata matrix chunks in the shared memory; According to the register mapping relationship matching the sparse computing unit, successively transfer at least one of the secondary chunks of each compressed data matrix, the secondary chunks of each index matrix, and the secondary chunks of each metadata matrix from the shared memory to the hardware registers.
15. The method according to claim 14, wherein, While locating the set of compressed matrices and the second associated matrix that match the first sparse matrix in the global memory of the computing chip, it further includes: Locating a selection array that matches the second associated matrix in the computing chip, where the selection array is used to describe the valid row positions in the second sparse matrix; the selection array of the second sparse matrix is the index of the non-zero data positions, and the selection array of the corresponding dense matrix of the second sparse matrix is the position mapping of the data in the dense matrix in the second sparse matrix. Correspondingly, transporting the second associated matrix from the global memory to the hardware register of the computing chip in the form of data blocks, including: Data-blocking the second associated matrix in the global memory according to the selection array. Successively transporting each data-block of the second associated matrix from the global memory to the shared memory. Performing secondary data-blocking on each data-block of the second associated matrix in the shared memory. According to the register mapping relationship matching the sparse computing unit, successively loading each data-block of the second associated matrix after secondary data-blocking from the shared memory to the hardware register.
16. The method according to claim 15, wherein After successively loading each data-block of the compressed data matrices, each data-block of the index matrices, and each data-block of the metadata matrices from the global memory to the shared memory, it further includes: Using a preset data rearrangement function to rearrange each data-block of the compressed data matrices, each data-block of the index matrices, and each data-block of the metadata matrices in the shared memory to avoid memory bank conflicts; and After successively transporting each data-block of the second associated matrix from the global memory to the shared memory, it further includes: Using the preset data rearrangement function to rearrange each data-block of the second associated matrix in the shared memory to avoid memory bank conflicts.
17. The method according to claim 15, wherein According to the register mapping relationship matching the sparse computing unit, successively transporting at least one of each data-block of the compressed data matrices after secondary data-blocking, each data-block of the index matrices after secondary data-blocking, and each data-block of the metadata matrices after secondary data-blocking from the shared memory to the hardware register, including: According to the register mapping relationship matching the sparse computing unit, calling a preset hardware instruction to successively transport each data-block of the compressed data matrices after secondary data-blocking and each data-block of the metadata matrices after secondary data-blocking, or each data-block of the compressed data matrices after secondary data-blocking, each data-block of the index matrices after secondary data-blocking, and each data-block of the metadata matrices after secondary data-blocking from the shared memory to the hardware register to achieve hardware acceleration of data loading; and According to the register mapping relationship matching the sparse computing unit, successively loading each data-block of the second associated matrix after secondary data-blocking from the shared memory to the hardware register, including: According to the register mapping relationship matching the sparse computing unit, calling a preset hardware instruction to successively load each data-block of the second associated matrix after secondary data-blocking from the shared memory to the hardware register to achieve hardware acceleration of data loading.
18. The method according to claim 15, wherein Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second correlation matrix according to the data loaded in batches in the hardware register, including: Through the sparse computing unit, obtain the current compressed data matrix quadratic block, the current metadata matrix quadratic block, and the current second correlation matrix quadratic block currently loaded in the hardware register; Through the sparse computing unit, generate a first sparse block matrix according to the current metadata matrix quadratic block and the current compressed data matrix quadratic block; Through the sparse computing unit, perform multiplication calculation according to the first sparse block matrix and the current second correlation matrix quadratic block.
19. The method according to claim 18, wherein Through the sparse computing unit, perform multiplication calculation according to the first sparse block matrix and the current second correlation matrix quadratic block, specifically including: Through the sparse computing unit, store the intermediate calculation amount generated during the multiplication calculation into the intermediate register of the computing chip; and Through the sparse computing unit, perform the matching multiplication calculation by obtaining the intermediate calculation amount from the intermediate register.
20. The method according to claim 19, wherein, If the second correlation matrix includes a second sparse matrix, after performing the multiplication calculation according to the first sparse block matrix and the current second correlation matrix quadratic block through the sparse computing unit, it further includes: After determining that the complete multiplication calculation result is stored in the intermediate register, convert the complete multiplication calculation result into a quadratic sparse block matrix according to the current index matrix quadratic block matching the current compressed data matrix quadratic block; Write back the quadratic sparse block matrix to the global memory of the computing chip level by level; If the second correlation matrix includes a dense matrix corresponding to the second sparse matrix, after performing the multiplication calculation according to the first sparse block matrix and the current second correlation matrix quadratic block through the sparse computing unit, it further includes: After determining that the complete multiplication calculation result is stored in the intermediate register, write back the multiplication calculation result to the global memory of the computing chip level by level in the original sparse format.
21. The method according to claim 13, wherein During each data block division process, V is an integer multiple of the division size in the column division direction; and In each sparse operator, the matrix specification of the first sparse matrix processed once is mi*ki, where V is an integer multiple of ki.
22. An optimization device for a mixture of experts model, characterized in that, It includes: A data structure sparse conversion module for performing sparse conversion on various sparse data structures of the target mixture-of-experts model to obtain a structured sparse data structure; wherein, the structured sparse data structure includes the structured sparse data structure of the model weight data, and the structured sparse data structure of the model weight data stores each sparse model weight matrix in the target mixture-of-experts model through a compressed matrix set; A structured sparse data structure optimization module for optimizing the data arrangement of the structured sparse data structure; A sparse operator optimization module for optimizing the operators of the target mixture-of-experts model in combination with the sparse computing unit of the computing chip to obtain sparse operators; A sparse operator update module, configured to update the original operator of the target mixture-of-experts model according to the sparse operator, so as to obtain a structured sparse mixture-of-experts model; A matrix set rearrangement storage module, configured to: obtain a sparse computing unit requirement file matching the computing chip; wherein, the sparse computing unit requirement file defines the data access positions of each thread executed within the sparse operator to each model weight matrix in the target mixture-of-experts model when each sparse computing unit on the computing chip performs sparse computing; rearrange the data arrangement modes of the compressed data matrix, the index matrix, and the metadata matrix in each of the compressed matrix sets according to the sparse computing unit requirement file; store the rearranged compressed matrix sets in a set memory of the computing chip, so as to realize continuous storage of data with continuous access characteristics.
23. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the optimization method of the mixture-of-experts model according to any one of claims 1-21.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the optimization method of the mixture-of-experts model according to any one of claims 1-21 when executed by a processor.
25. A computer program product comprising a computer program / instructions, wherein, The computer program / instructions, when executed by a processor, implement the optimization method of the mixture-of-experts model according to any one of claims 1-21.