Operator Optimization Method, Device, Equipment, Medium and Program for Mixture-of-Experts Model
By using data chunking and sparse calculation units to generate sparse operators in the hybrid expert model, the problem of neglected sparseness of input data in the hybrid expert model is solved, and operator optimization and computational performance improvement of the hybrid expert model is achieved.
Patent Information
- Application Number
- CN202510308464.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing hybrid expert model computing technology mainly focuses on solving parameter redundancy problems, ignores the sparse characteristics of input data, and lacks hardware instruction support to provide structured sparse computing acceleration.
By combining the sparse calculation unit of the computing chip using data chunking, multiple sparse operators suitable for the original hybrid expert model are generated, and the original model is updated based on these sparse operators to obtain the target mixed expert model.
In the model computing scenario of hybrid expert models, the hardware acceleration performance of sparse computing units is fully utilized to optimize the calculation, bandwidth and storage resource overhead of sparse matrix multiplication operations.
Smart Images

Figure CN119830976B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to an operator optimization method, device, electronic device, storage medium, and program for a mixture of experts model. Background Art
[0002] A mixture of experts model (MoE) improves the generalization ability of a model by combining multiple expert models to process different input data, and has currently been widely integrated into large language models.
[0003] The sparse activation characteristic of the mixture of experts model poses new challenges to storage and computing resources. On the one hand, the parameters of large language models have a certain degree of redundancy, and a sparse method can be used to reduce the amount of computation while ensuring the model accuracy; on the other hand, due to the existence of an expert selection strategy in the mixture of experts model, the sparsity of input data is introduced.
[0004] In the process of implementing the present invention, the inventors found that current mixture of experts computing technologies mainly focus on using sparse computing technologies to solve the redundancy problem of parameters, while ignoring the sparse characteristics of input data. In addition, there is a lack of using hardware instruction support to provide structured sparse computing acceleration in mixture of experts computing. Summary of the Invention
[0005] Embodiments of the present invention provide an operator optimization method, device, electronic device, storage medium, and program for a mixture of experts model, to achieve operator optimization for the mixture of experts model, give full play to the hardware acceleration performance of sparse computing units in a computing chip in the model computing scenario of the mixture of experts model, and thus greatly optimize the computation, bandwidth, and storage resource overheads in the sparse matrix multiplication operation process.
[0006] According to one aspect of the present invention, there is provided an operator optimization method for a mixture of experts model, including:
[0007] Combining a sparse computing unit of a computing chip in a data chunking manner to generate sparse operators applicable to an original mixture of experts model; wherein the number of the sparse operators is multiple;
[0008] Updating original operators of the original mixture of experts model according to the sparse operators to obtain a target mixture of experts model.
[0009] According to another aspect of the present invention, there is provided an operator optimization device for a mixture of experts model, including:
[0010] A sparse operator generation module, configured to combine a sparse computing unit of a computing chip in a data chunking manner to generate sparse operators applicable to an original mixture of experts model; wherein the number of the sparse operators is multiple;
[0011] A sparse operator update module, configured to update the original operator of the original mixture-of-experts model according to the sparse operator, so as to obtain a target mixture-of-experts model.
[0012] According to another aspect of the present invention, there is provided an electronic device, including:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the operator optimization method of the mixture-of-experts model according to any embodiment of the present invention.
[0016] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a processor, the operator optimization method of the mixture-of-experts model according to any embodiment of the present invention is implemented.
[0017] According to another aspect of the present invention, there is also provided a computer program product including a computer program, and when the computer program is executed by a processor, the operator optimization method of the mixture-of-experts model according to any embodiment of the present invention is implemented.
[0018] In the embodiments of the present invention, by adopting a data chunking method in combination with the sparse computing unit of the computing chip, a plurality of sparse operators suitable for the original mixture-of-experts model are generated, so as to update the original operator of the original mixture-of-experts model according to the generated sparse operators, and a target mixture-of-experts model is obtained. The sparse characteristics during matrix multiplication calculation of the mixture-of-experts model based on the sparse matrix are fully considered. By adopting a data chunking and data transfer technology adapted to the sparse characteristics, the operator optimization of the mixture-of-experts model is realized, and the hardware acceleration performance of the sparse computing unit in the computing chip is fully exerted in the model calculation scenario of the mixture-of-experts model, thereby greatly optimizing the computing, bandwidth, and storage resource overheads during the sparse matrix multiplication operation.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of an operator optimization method for a mixture of experts model provided by an embodiment of the present invention;
[0022] Figure 2 is a flowchart of another operator optimization method for a mixture of experts model provided by an embodiment of the present invention;
[0023] Figure 3 is a schematic flowchart of optimizing the original operator of a mixture of experts model provided by an embodiment of the present invention;
[0024] Figure 4 is a flowchart of a method for accelerating the multiplication of sparse matrices provided by an embodiment of the present invention;
[0025] Figure 5 is a schematic diagram of the comparison structure between a first sparse matrix and a matching set of compressed matrices applicable to an embodiment of the present invention;
[0026] Figure 6 is a block diagram of the implementation of data chunking and multi-level transfer processes applicable to an embodiment of the present invention;
[0027] Figure 7 is a schematic diagram of a register mapping relationship applicable to an embodiment of the present invention;
[0028] Figure 8 is a schematic diagram of the structure of a sparse computing unit cooperating with an intermediate register to perform calculations applicable to an embodiment of the present invention;
[0029] Figure 9 is a schematic diagram of an operator optimization device for a mixture of experts model provided by an embodiment of the present invention;
[0030] Figure 10 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0031] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] Figure 1 is a flowchart of an operator optimization method for a mixture of experts model provided by an embodiment of the present invention. This embodiment is applicable to the situation where the original operator of the mixture of experts model is optimized and updated by a sparse operator obtained by optimizing in a data chunking manner. This method can be executed by an operator optimization device of the mixture of experts model. The device can be implemented in a software and / or hardware manner and is generally integrated in an electronic device. The electronic device can be a terminal device or a server device, as long as it can execute the operator optimization method of the mixture of experts model. The specific device type of the electronic device is not limited in the embodiments of the present invention. Correspondingly, as Figure 1 shown, the method includes the following operations:
[0034] S110. Generate sparse operators applicable to the original mixture of experts model by combining the sparse computing units of the computing chip in a data chunking manner; wherein, the number of the sparse operators is multiple.
[0035] Among them, the sparse computing unit can be a hardware computing unit in the computing chip. The original mixture of experts model can be a mixture of experts model that needs to optimize the internal operators. The model type and model structure of the original mixture of experts model are not limited in the embodiments of the present invention.
[0036] Optionally, a computing chip can be understood as an integrated circuit used to implement a set computing task (for example, Internet of Things control, high-performance computing, or mobile computing, etc.). The computing chip can be a general-purpose computing chip, such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a heterogeneous chip combination of "CPU + GPU equipped with a sparse acceleration computing unit", or a dedicated computing chip, such as an NPU (Neural Processing Unit). The embodiments of the present invention do not limit this.
[0037] Among them, a sparse operator refers to an operator that only operates on some elements, such as a sparse matrix in matrix multiplication. During the calculation process, the processing of sparse operators usually involves how to effectively store and calculate sparse matrices, and how to optimize the calculation performance of sparse operators. Sparse operators are mainly applied to process sparse matrices, in which most elements are zero and only a small number of elements are non-zero. Sparse operators can reduce the use of storage space through compression storage methods, accelerate operations such as sparse matrix multiplication using algorithms such as the fast Fourier transform, and at the same time use memory optimization techniques (such as cache optimization and memory alignment) to improve memory usage efficiency. They can also utilize parallel computing techniques (such as multi-threading or distributed computing) to accelerate the operations of large-scale sparse matrices.
[0038] In the embodiments of the present invention, in order to achieve the computing acceleration of adapting to the sparse acceleration hardware computing unit, fully considering the sparse characteristics of the operation data of the sparse operator, a data chunking method can be adopted to combine with the sparse computing unit of the computing chip to optimize the computing, bandwidth, and storage resource overhead during the operator execution process, and through compilation optimization means such as operator fusion, generate various types of sparse operators applicable to the original mixture-of-experts model.
[0039] S120. Update the original operator of the original mixture-of-experts model according to the sparse operator to obtain a target mixture-of-experts model.
[0040] Among them, the original operator can be the original operator type of the original mixture-of-experts model. The target mixture-of-experts model can be a mixture-of-experts model with an optimized structure obtained by optimizing and updating the original operator of the original mixture-of-experts model according to the sparse operator.
[0041] It can be understood that the original operator of the original mixture-of-experts model is usually a matrix multiplication operator. The matrix multiplication operator usually performs a dense matrix multiplication operation.
[0042] Correspondingly, after generating various types of sparse operators applicable to the original mixture-of-experts model, the types and functions of the original operators of the original mixture-of-experts model can be considered, and the original operators are updated and replaced with adapted sparse operators to achieve operator optimization and structural update of the original mixture-of-experts model, and a target mixture-of-experts model that can be accelerated by hardware instructions is obtained.
[0043] Correspondingly, the optimized target mixture-of-experts model can take a structured sparse matrix as input, optimize the data flow of the sparse operator in this format, and at the same time arrange the calculation process of the computing cores, and use the sparse computing units in the computing chip at the hardware level to achieve optimization and acceleration of the sparse operator, and finally inject it into the process of mixture-of-experts calculation, thereby accelerating the execution of the mixture-of-experts model.
[0044] It can be seen that through the operator optimization method of the mixture-of-experts model provided by the embodiments of the present invention, a structured sparse conversion applicable to mixture-of-experts calculation can be realized. Finally, the converted model replaces the original mixture-of-experts model, and the model calculation is performed in the same way as before, and the calculation acceleration is achieved by calling a specific sparse acceleration unit.
[0045] The embodiments of the present invention generate multiple sparse operators applicable to the original mixture-of-experts model by adopting a data chunking method in combination with the sparse computing units of the computing chip, and thus update the original operators of the original mixture-of-experts model according to the generated sparse operators to obtain a target mixture-of-experts model. The sparse characteristics during matrix multiplication calculation based on the sparse matrix in the mixture-of-experts model are fully considered. By adopting data chunking and data transfer techniques adapted to the sparse characteristics, the operator optimization of the mixture-of-experts model is realized, and the hardware acceleration performance of the sparse computing units in the computing chip is fully exerted in the model calculation scenario of the mixture-of-experts model, thereby greatly optimizing the overhead of calculation, bandwidth, and storage resources during the sparse matrix multiplication operation.
[0046] Figure 2 is a flowchart of another operator optimization method of the mixture-of-experts model provided by the embodiments of the present invention. Figure 3 is a schematic flowchart of optimizing the original operators of the mixture-of-experts model provided by the embodiments of the present invention. This embodiment is specific based on the above embodiment. In this embodiment, various specific optional generation methods of different types of sparse operators are given. Correspondingly, as Figure 2 and Figure 3 shown, the method of this embodiment may include:
[0047] S210. Adopt a data chunking method in combination with the sparse computing units of the computing chip to generate a first sparse operator, a second sparse operator, and a third sparse operator applicable to the original mixture-of-experts model.
[0048] Among them, the first sparse operator may be an operator used to output the calculation result of the model in the mixture of experts model. The second sparse operator may be an operator capable of outputting the output value of the activation function. The third sparse operator may be an operator capable of outputting the output value of the dot product function.
[0049] In an alternative embodiment of the present invention, if the sparse operator includes the first sparse operator, the method of combining the sparse computing unit of the computing chip by means of data block division to generate a sparse operator applicable to the original mixture of experts model may include: locating a set of compressed matrices and a second correlation matrix matching the first sparse matrix in the global memory of the computing chip according to the first sparse matrix multiplication requirement; wherein, the second correlation matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; wherein, the first sparse matrix includes a plurality of compression units of a set size, each of the compression units includes at least one valid row and at least one sparse row, each of the valid rows includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of the valid rows in the first sparse matrix in the corresponding compression units, and a metadata matrix for storing the positions of the non-zero data in the corresponding structured sparse computing storage units; transporting the set of compressed matrices and the second correlation matrix to the hardware registers of the computing chip step by step in the form of data blocks from the global memory; through the sparse computing unit of the computing chip, gradually calculating the multiplication result of the first sparse matrix and the second correlation matrix according to the data loaded in batches in the hardware registers; processing the multiplication result of the first sparse matrix and the second correlation matrix according to the processing logic of the broadcast multiplication function of the original mixture of experts model to obtain the first sparse operator.
[0050] In an alternative embodiment of the present invention, if the sparse operator includes a second sparse operator, then generating a sparse operator applicable to the original mixture-of-experts model by combining the sparse computing units of the computing chip in a data-blocking manner may include: locating, according to the second sparse matrix multiplication requirement, a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; wherein, the first sparse matrix includes a plurality of compressed units of a set size, each of the compressed units includes at least one valid row and at least one sparse row, each of the valid rows includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of the valid rows in the first sparse matrix in the respective compressed units, and a metadata matrix for storing the positions of the non-zero data in the respective structured sparse computing storage units; transporting the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data-blocking form, and during the transportation, rearranging or transforming the data in the original global memory; through the sparse computing units of the computing chip, gradually calculating the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; processing the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the activation function of the original mixture-of-experts model to obtain the second sparse operator.
[0051] In an alternative embodiment of the present invention, if the sparse operator includes a third sparse operator, then generating a sparse operator applicable to the original mixture-of-experts model by combining the sparse computing units of the computing chip in a data-blocking manner may include: locating, according to the third sparse matrix multiplication requirement, a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; wherein, the first sparse matrix includes a plurality of compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing storage unit; transporting the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data-blocking form, and during the transportation, rearranging or deforming the data in the original global memory; through the sparse computing units of the computing chip, gradually calculating the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; processing the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the dot product function of the original mixture-of-experts model to obtain the third sparse operator.
[0052] Among them, the first sparse matrix can be the left operand of a double matrix multiplication and is a sparse matrix. The second associated matrix can be the right operand of a double matrix multiplication and can be a sparse matrix or a dense matrix composed of valid data in the sparse matrix. The matrix multiplication between the first sparse matrix and the second associated matrix can be referred to as sparse matrix multiplication. When the second associated matrix also uses a sparse matrix, the matrix multiplication between the first sparse matrix and the second associated matrix can be referred to as double sparse matrix multiplication. In the following text, regardless of the type of matrix used for the second associated matrix, the matrix multiplication between the first sparse matrix and the second associated matrix is simply referred to as sparse matrix multiplication.
[0053] In the embodiments of the present invention, the above-mentioned sparse matrix multiplication acceleration method can be encapsulated in the form of an operator interface. For example, various types of sparse operators for implementing matrix multiplication between a first sparse matrix and a second associated matrix are specifically constructed, such as a first sparse operator, a second sparse operator, or a third sparse operator, etc. Furthermore, by calling the operator interface of the sparse operator, the multiplication acceleration method between the first sparse matrix and the second associated matrix can be triggered to execute, and additional processing logic can be configured for the multiplication result between the first sparse matrix and the second associated matrix according to the fusion requirements of the first sparse operator and other functional functions, thereby generating various types of sparse operators, such as a first sparse operator, a second sparse operator, or a third sparse operator, etc.
[0054] Specifically, sparse matrix multiplication can be understood as that at least one type of operand in the left operand and the right operand for performing matrix multiplication is a sparse matrix. Correspondingly, when a first sparse operator is generated upon detecting an interface call request for the first sparse operator, it is determined that a first sparse matrix multiplication requirement is detected; when a second sparse operator is generated upon detecting an interface call request for the second sparse operator, it is determined that a second sparse matrix multiplication requirement is detected; when a third sparse operator is generated upon detecting an interface call request for the third sparse operator, it is determined that a third sparse matrix multiplication requirement is detected. Furthermore, the identification information of the left operand (hereinafter referred to as the first sparse matrix) and the right operand (hereinafter referred to as the second associated matrix) that need to perform the multiplication of the sparse matrix can be obtained from the corresponding interface call request. Based on the above identification information, the compressed matrix set matching the first sparse matrix and the second associated matrix can be located in the global memory of the computing chip.
[0055] It can be understood that the first sparse matrix and the second associated matrix perform multiplication calculations in the computing chip. Furthermore, the first sparse matrix and the second associated matrix that need to be calculated are pre-loaded into the computing chip. To improve the computing speed of the computing chip, the computing data is initially stored in the global memory of the computing chip, and subsequently, the computing data can be transferred in blocks to the hardware register adapted to the hardware computing unit in a step-by-step manner, and the final computing result is obtained by the hardware computing unit based on the data blocks stored in the hardware register.
[0056] In this embodiment, in order to achieve the final multiplication acceleration of the sparse operator, a special limitation is imposed on the sparse format of the first sparse matrix as the left operand. At the same time, the storage method of the first sparse matrix in the global memory of the computing chip is also optimized accordingly.
[0057] Among them, the first sparse matrix includes a plurality of compression units of a set size. Each compression unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse pattern; the compression matrix set includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of the valid rows in the first sparse matrix within their respective compression units, and a metadata matrix for storing the positions of the non-zero data within their respective structured sparse computing storage units.
[0058] Specifically, the first sparse matrix includes an integer number (one or more) of compression units. That is, the size (number of rows * number of columns) of the first sparse matrix is divisible by the size of the compression unit. A compression unit can be understood as a small matrix of a specific size, and this small matrix has at least two matrix rows. Among the above at least two matrix rows, there is at least one valid row and at least one sparse row. A valid row can be understood as a matrix row that includes at least one valid data (non-zero), and a sparse row refers to a matrix row where all data are 0.
[0059] Furthermore, each valid row in the compression unit includes an integer number (one or more) of structured sparse computing storage units. Among them, all the structured sparse computing storage units included in the first sparse matrix are in the same sparse pattern. Generally speaking, for the convenience of calculation, the size of the structured sparse computing storage unit can be adapted to the computing scale of the hardware computing unit in the computing chip.
[0060] The sparse pattern can be understood as the proportion of valid data (non-zero value data) in all data. For example, the sparse pattern can be 2:4 or 4:8, etc. That is, for a 2:4 sparse pattern, the structured sparse computing storage unit altogether includes 4 data, and 2 of these data are non-zero data. And the arrangement positions of the above two non-zero data in the structured sparse computing storage unit are not restricted.
[0061] It can be understood that in the first sparse matrix, there are a large number of 0 value data. If the first sparse matrix is directly stored in the global memory of the computing chip according to the original size of the first sparse matrix, it will cause a waste of a large number of storage units. In addition, directly implementing subsequent sparse matrix multiplication based on the first sparse matrix of the original size has low computing efficiency and cannot efficiently use the hardware computing unit in the computing chip, that is, the sparse computing unit.
[0062] In view of this, in the embodiments of the present invention, for the left operand (the first sparse matrix) of the sparse multiplication calculation with the above special structure, a novel and efficient data compression method is creatively proposed to solve the above various technical problems.
[0063] In an embodiment of the present invention, through a specific data compression method, the first sparse matrix can be tightly stored in the form of a set of compressed matrices. Specifically, three compressed matrices corresponding to the first sparse matrix are stored in the set of compressed matrices, namely a compressed data matrix, an index matrix, and a metadata matrix.
[0064] Among them, the compressed data matrix is used to store the compressed data of the non-zero data in the first sparse matrix, the index matrix is used to store the positions of the valid rows in the first sparse matrix in the corresponding compression unit, and the metadata matrix is used to store the positions of the non-zero data in the corresponding structured sparse computing storage unit.
[0065] Obviously, through the above data compression method, the specific positions of each non-zero data in the original first sparse matrix can be determined by three small matrices. This data compression method can also effectively reduce the consumption of storage resources in the computing chip and facilitate calculations based on the first sparse matrix.
[0066] In an embodiment of the present invention, the second correlation matrix may not be compressed and stored. However, considering the actual application requirements of sparse matrix multiplication, for example, when applied in a mixture of experts model, the second correlation matrix generally also has a specific data structure. Typically, when the second correlation matrix is a second sparse matrix, one or more matrix rows in the second sparse matrix are sparse rows (all 0). When the second correlation matrix is a dense matrix corresponding to the second sparse matrix, the data in the dense matrix corresponding to the second sparse matrix is the valid data in the second sparse matrix. Exemplarily, assuming the second sparse matrix is [1, 3, 0, 5, 0, 0], the dense matrix corresponding to the second sparse matrix is [1, 3, 5].
[0067] Based on this, in the global memory of the computing chip, in addition to storing the second correlation matrix in its entirety, a selection array matching the second correlation matrix is further stored. The selection array is used to describe the positions of the valid rows in the second correlation matrix. In a specific example, assuming the second correlation matrix is a second sparse matrix, and all rows except the 0th, 3rd, and 5th rows in the second sparse matrix are sparse rows, a selection array in the form of {0, 3, 5} can be constructed and stored in association with the second correlation matrix. Assuming the second correlation matrix is a dense matrix corresponding to the second sparse matrix, and the second sparse matrix includes the valid data in rows 1 - 3 and does not include sparse rows, a selection array in the form of {1, 2, 3} can be constructed and stored in association with the second correlation matrix.
[0068] Correspondingly, when locating the second correlation matrix in the global memory of the computing chip, the selection array matching the second correlation matrix can be located synchronously.
[0069] In an alternative embodiment of the present invention, the computing chip includes a three-level storage architecture of global memory -> shared memory -> hardware registers. Among them, the above three-level storage is physically closer and closer to the hardware computing units used for computing in the computing chip. In particular, the hardware registers are arranged close to the hardware computing units.
[0070] Generally speaking, the closer the storage device is to the computing unit, the faster its data reading and writing speed, but its storage capacity is also smaller. Therefore, for a matrix multiplication of C(M×N)=A(M×K)×B(K×N), the complete data of A and B cannot be completely stored in the shared memory or hardware registers. Often, it is necessary to transfer them in chunks through a multi-level transfer method to the hardware registers, so that the hardware computing units in the computing chip can calculate the final complete multiplication result in multiple steps.
[0071] In an alternative embodiment of this embodiment, the data of each matrix in the compressed matrix set can be processed in chunks by combining the special data structure of the compressed matrix set. At the same time, the data of the second associated matrix can be processed in chunks by combining the selection array matching the second sparse matrix, so as to improve the data transfer and subsequent multiplication calculation efficiency.
[0072] In this embodiment, the sparse computing unit can be understood as a special hardware circuit in the computing chip for performing matrix multiplication on sparse matrices. Such special hardware circuits can accelerate sparse matrix multiplication. Generally, there are multiple sparse computing units in the computing chip. Inside each sparse computing unit, one or more threads can be started to perform corresponding sparse computing tasks.
[0073] By using the cooperation of each sparse computing unit in the computing chip, partial data in the first sparse matrix and the second associated matrix can be loaded in multiple steps from each hardware register for local calculation, and the local calculation results can be combined or accumulated to finally obtain the multiplication result of the first sparse matrix and the second associated matrix.
[0074] In the calculation process of the mixture of experts model, there are many one-to-one operators adjacent to matrix multiplication, but the calculation of two consecutive operators will introduce redundant calculation kernel startup overhead and I / O (Input / Output) overhead. In order to eliminate the overhead of loading and storing intermediate results and multiple kernel startups, the calculation cores of each sparse operator can be improved by using the method of operator fusion.
[0075] Among various sparse operators, the first sparse operator is usually the operator for outputting the calculation result of the model. Exemplarily, such as Figure 3As shown, the sparse operator 3 in the target mixture-of-experts model can be used as the first sparse operator. Optionally, for the first sparse operator, after the multiplication result of the first sparse matrix and the second correlation matrix is gradually calculated by the sparse computing unit of the computing chip according to the data loaded in batches in the hardware register, the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the broadcast multiplication function of the original mixture-of-experts model to obtain the first sparse operator with broadcast function.
[0076] As Figure 3 shown, the sparse operator 1 in the target mixture-of-experts model can be used as the second sparse operator. Optionally, for the second sparse operator, after the multiplication result of the first sparse matrix and the second correlation matrix is gradually calculated by the sparse computing unit of the computing chip according to the data loaded in batches in the hardware register, the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the activation function of the original mixture-of-experts model to obtain the second sparse operator. Correspondingly, the second sparse operator can directly output the calculation result of the activation function.
[0077] As Figure 3 shown, the sparse operator 2 in the target mixture-of-experts model can be used as the third sparse operator. Optionally, for the third sparse operator, after the multiplication result of the first sparse matrix and the second correlation matrix is gradually calculated by the sparse computing unit of the computing chip according to the data loaded in batches in the hardware register, the multiplication result of the first sparse matrix and the second correlation matrix can be processed according to the processing logic of the dot product function of the original mixture-of-experts model to obtain the third sparse operator. Correspondingly, the third sparse operator can directly output the calculation result of the dot product function.
[0078] For the first sparse operator, the second sparse operator, and the third sparse operator, the generation process of the overall sparse operator is generally the same. The difference is only that in the process of gradually transferring data from the global memory to the hardware register of the computing chip in the form of data blocks, the second sparse operator and the third sparse operator need to reorder or transform the data in the original global memory during the transfer process. At the same time, because the operator functions are different, the types of specific functional functions fused by the first sparse operator, the second sparse operator, and the third sparse operator during operator fusion are different.
[0079] S220. Construct a sparse operator library according to the first sparse operator, the second sparse operator, and the third sparse operator.
[0080] Among them, the sparse operator library can be used to store various generated sparse operators.
[0081] In the embodiments of the present invention, as Figure 3As shown, after generating various sparse operators suitable for the original mixture-of-experts model by using the data chunking method in combination with the sparse computing unit of the computing chip, the generated various sparse operators can be stored in the sparse operator library. It can be understood that the above process of generating sparse operators only schematically illustrates the specific generation processes of the first sparse operator, the second sparse operator, and the third sparse operator. For other types of sparse operators in the original mixture-of-experts model, similar methods can also be used for optimized generation. That is to say, the sparse operators are not limited to the first sparse operator, the second sparse operator, and the third sparse operator. Correspondingly, in addition to storing the first sparse operator, the second sparse operator, and the third sparse operator, the sparse operator library can also store other types of sparse operators, and the embodiments of the present invention do not limit the types of sparse operators.
[0082] S230. Screen target sparse operators from the sparse operator library according to the type of the original mixture-of-experts model, and update the original operators of the original mixture-of-experts model according to the target sparse operators to obtain a target mixture-of-experts model.
[0083] Among them, the target sparse operator can be a sparse operator used to update and replace the original operators in the original mixture-of-experts model. Optionally, the number of target sparse operators can be multiple.
[0084] Exemplarily, as Figure 3 shown, the original operators of the original mixture-of-experts model can include a first matrix multiplication operator, a second matrix multiplication operator, and a third matrix multiplication operator. Among them, the first matrix multiplication operator can be the matrix multiplication 2 operator in the original mixture-of-experts model, the second matrix multiplication operator can be the matrix multiplication 1 operator in the original mixture-of-experts model, and the third matrix multiplication operator can be the matrix multiplication 3 operator in the original mixture-of-experts model. For this type of original mixture-of-experts model, the first sparse operator, the second sparse operator, and the third sparse operator can be screened as the target sparse operators. Correspondingly, when updating the original operators of the original mixture-of-experts model according to the target sparse operators, the first matrix multiplication operator of the original mixture-of-experts model can be replaced and updated according to the first sparse operator, the second matrix multiplication operator of the original mixture-of-experts model can be replaced and updated according to the second sparse operator, and the third matrix multiplication operator of the original mixture-of-experts model can be replaced and updated according to the third sparse operator to obtain a target mixture-of-experts model.
[0085] Exemplarily, as Figure 3As shown, the original operators of some types of original mixture-of-experts models only include a first matrix multiplication operator and a second matrix multiplication operator. Among them, the first matrix multiplication operator can be the matrix multiplication 2 operator in the original mixture-of-experts model, and the second matrix multiplication operator can be the matrix multiplication 1 operator in the original mixture-of-experts model. For this type of original mixture-of-experts model, a first sparse operator and a second sparse operator can be selected as target sparse operators. Correspondingly, when updating the original operators of the original mixture-of-experts model according to the target sparse operators, the first matrix multiplication operator of the original mixture-of-experts model can be replaced and updated according to the first sparse operator, and the second matrix multiplication operator of the original mixture-of-experts model can be replaced and updated according to the second sparse operator to obtain a target mixture-of-experts model.
[0086] Similarly, for other types of original mixture-of-experts models, sparse operator types with other functions can also be adaptively generated, and the original operators inside can be replaced and updated to obtain corresponding optimized target mixture-of-experts models.
[0087] The technical solution of the embodiment of the present invention, after compressing and storing the sparse matrix that needs to perform multiplication calculations in a specific compression format, transfers the sparse matrix in the above specific compression format and the associated matrix to the hardware registers of the computing chip level by level in the form of data blocks from the global memory, and through the sparse computing unit of the computing chip, according to the data loaded in batches in the hardware registers, gradually calculates the multiplication result of the sparse matrix and the associated matrix. This implementation method fully considers the sparse characteristics when performing matrix multiplication based on the sparse matrix. By adopting data block and data transfer technologies adapted to this sparse characteristic, the hardware acceleration performance of the sparse computing unit in the computing chip can be fully utilized, and thus the overhead of computing, bandwidth, and storage resources during the sparse matrix multiplication operation can be greatly optimized, which is particularly suitable for the model calculation scenario of the mixture-of-experts model.
[0088] Figure 4 It is a flowchart of a method for accelerating the multiplication of sparse matrices provided by an embodiment of the present invention. This embodiment is optimized based on the above embodiments. In this embodiment, the operation of "transferring the compressed matrix set and the second associated matrix to the hardware registers of the computing chip level by level in the form of data blocks" is specifically implemented.
[0089] Correspondingly, as Figure 4 shown, the method may include:
[0090] S410. According to the sparse matrix multiplication requirement, locate the compressed matrix set and the second associated matrix that match the first sparse matrix in the global memory of the computing chip.
[0091] Among them, the sparse matrix multiplication requirements may include, but are not limited to, the first sparse matrix multiplication requirement, the second sparse matrix multiplication requirement, the third sparse matrix multiplication requirement, etc.
[0092] The inventors found through research that when using a mixture of experts model to calculate tasks (typically, model inference tasks), the matrix multiplication calculation between the model weight matrix of each model layer and the sparse matrix of intermediate activation values input to the mixture model layer in the mixture of experts model belongs to double sparse matrix multiplication. The matrix multiplication calculation between the model weight matrix of each model layer and the dense matrix of intermediate activation values input to the mixture model layer in the mixture of experts model belongs to sparse matrix multiplication. Furthermore, the methods of the embodiments of the present invention can be applied to the model calculation and model optimization scenarios of the mixture of experts model.
[0093] In the calculation scenario of the mixture of experts model, by performing a sparse matrix multiplication process based on a specific model weight matrix and a sparse matrix or dense matrix of matching intermediate activation values, and combining an operator fusion strategy, various types of sparse operators can be generated to achieve operator optimization of the mixture of experts model, and further complete the optimization of the internal model structure of the mixture of experts model.
[0094] Correspondingly, in an optional implementation manner of the embodiments of the present invention, the first sparse matrix is the model weight matrix in the target mixture of experts model, and the second associated matrix is the sparse matrix or dense matrix of intermediate activation values input to the mixture model layer in the target mixture of experts model.
[0095] Furthermore, the first sparse matrix may specifically include multiple compression units of M*V size;
[0096] Each compression unit specifically includes d valid rows, where 1≤d<M; and
[0097] Each structured sparse calculation and storage unit is in an N:L sparse mode, where L is the total number of data included in the structured sparse calculation and storage unit, and N is the number of non-zero data included in the structured sparse calculation and storage unit.
[0098] Correspondingly, for a first sparse matrix of m*k size, a compressed data matrix of (m / M*d)*(k / (L / N)) size, an index matrix of (m / M*d)*(k / V) size, and a metadata matrix of (m / M*d)*(k / (L / N)) size can be obtained.
[0099] For the sake of convenience of explanation, Figure 5 shows a schematic diagram of the control structure of a first sparse matrix and a set of matching compressed matrices applicable to the embodiments of the present invention.
[0100] Specifically, as Figure 5The diagram shown is a schematic diagram of a set of compressed matrices matching a first sparse matrix with m=4, k=16. In the first sparse matrix, there are 4 compression units with M=2, V=8, each of which contains d=1 valid rows, each of which contains 2 structured sparse computing storage units, and each structured sparse computing storage unit is in a sparse mode with N:L of 2:4. That is, A, B, ..., L filled in each matrix position of the first sparse matrix represent non-zero data, and the blank position represents the data filled with zero value.
[0101] Correspondingly, the size of the compressed data matrix corresponding to the first sparse matrix is (m / M*d)*(k / (L / N))=2*8. The compressed data matrix stores the non-zero data A, B, ..., L in the first sparse matrix in sequence.
[0102] Furthermore, the size of the index matrix corresponding to the first sparse matrix is (m / M*d)*(k / V)=2*2, and the positions of each valid row in the first sparse matrix in the corresponding compression unit are stored in different positions of the index matrix. Figure 5 As shown, the first sparse matrix includes a 2*8 compression unit 1 where A, B, C, and D are located. Since d=1, and the row position where the valid row in the compression unit 1 is located, that is, the row where A, B, C, and D are located, is the 0th row, and further, the above positional relationship can be identified by the 0 stored at the matrix position in the upper left corner of the index matrix.
[0103] Furthermore, the size of the metadata matrix corresponding to the first sparse matrix is (m / M*d)*(k / (L / N))=2*8, and the positions of each non-zero data in the corresponding structured sparse computing storage unit are stored in different positions of the metadata matrix. Figure 5 As shown, the structured sparse computing storage unit 1 containing non-zero elements A and B is located in the upper left corner of the first sparse matrix. Furthermore, the 0 and 2 at the first two column positions of the first row of the metadata matrix represent the column positions of A and B in the structured sparse computing storage unit 1, respectively, with A being located in the 0th column and B being located in the 2nd column.
[0104] It is understandable that by constructing the above-mentioned compressed data matrix, index matrix and metadata data, the first sparse matrix of the above-mentioned specific structure can be uniquely determined. Furthermore, the sparse model weight matrices contained in the target hybrid expert model can be efficiently compressed and stored, thereby effectively reducing the storage overhead of the target hybrid expert model on the configured computing chip.
[0105] S420. Chunk the compressed data matrix in the global memory, and perform chunking of the index matrix and the metadata matrix in a manner that matches the data correspondence among the compressed data matrix, the index matrix, and the metadata matrix.
[0106] In the embodiments of the present invention, since the compressed data matrix stores all non-zero data in the first sparse matrix, data chunking can be performed according to the data scale of the compressed data matrix. During the process of chunking the compressed data matrix, to further consider the efficiency of subsequent data transfer and multiplication calculations, it can be set that during the process of chunking the compressed data matrix in the global memory, V (the number of columns included in each compression unit) is an integer multiple of the chunking size in the column splitting direction.
[0107] For example, assume that in the global memory, when chunking the compressed data matrix into multiple chunks of size m b *k b it is required that V is an integer multiple of k b For example, k b is V / 2 or V / 4, etc.
[0108] As shown above, since there is a data or position correspondence among the compressed data matrix, the index matrix, and the metadata matrix. After determining the chunks of the compressed data matrix sliced from the compressed data matrix, the corresponding chunks of the index matrix and the metadata matrix that match the chunks of the compressed data matrix can be uniquely determined. That is, perform chunking of the index matrix and the metadata matrix in a manner that matches the data correspondence among the compressed data matrix, the index matrix, and the metadata matrix.
[0109] S430. Locate the selection array that matches the second associated matrix in the computing chip, and according to the selection array, perform data chunking on the second associated matrix in the global memory.
[0110] Among them, the selection array of the second sparse matrix is the index of the positions of non-zero data therein, and the selection array of the dense matrix corresponding to the second sparse matrix is the position mapping of the data in the dense matrix in the second sparse matrix.
[0111] As shown above, if the second associated matrix is a second sparse matrix, the second sparse matrix may contain a large number of sparse rows. If these sparse rows are also chunked and sent to the hardware computing unit for matrix multiplication calculation, it will bring redundant data loading and calculation, reducing the calculation efficiency.
[0112] Based on this, in the embodiments of the present invention, in the global memory, data chunking can be performed only on the valid data rows in the second association matrix based on the valid row positions defined by the selection array. Optionally, one or more valid data rows can be taken as a data chunk each time, or a set number of data in one valid data row can be taken as a data chunk each time, etc.
[0113] S440. Successively load each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk from the global memory into the shared memory.
[0114] After data chunking for the set of compressed matrices is completed in the global memory, each of the above-mentioned compressed data matrix chunks, each index matrix chunk, and each metadata matrix chunk can be successively loaded from the global memory into the shared memory.
[0115] Based on the above embodiments, in order to prevent the data in the shared memory from being restricted by the bank conflict where the same memory bank in the computing chip cannot be accessed by multiple threads simultaneously during access (read / write), in the embodiments of the present invention, data rearrangement in the shared memory is further considered using a specific data rearrangement function based on the data offset.
[0116] Correspondingly, in an optional implementation manner of the embodiments of the present invention, after successively transferring each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk from the global memory to the shared memory, it may further include:
[0117] Use a preset data rearrangement function to rearrange each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk in the shared memory to avoid bank conflicts.
[0118] Specifically, according to the specific type and model of the computing chip, a matching data rearrangement function can be selected from the adapted function library to rearrange the data in the shared memory, and the embodiments of the present invention do not limit this.
[0119] S450. Successively transfer each second association matrix chunk from the global memory to the shared memory.
[0120] Similarly, after data chunking of the second association matrix is completed in the global memory, each second association matrix chunk can be correspondingly successively transferred from the global memory to the shared memory.
[0121] Meanwhile, in order to avoid the problem of bank conflicts, in an optional implementation manner of the embodiments of the present invention, after successively transferring each second association matrix chunk from the global memory to the shared memory, it may further include:
[0122] Use a preset data rearrangement function to rearrange each second correlation matrix block in shared memory to avoid bank conflicts.
[0123] Of course, it can be understood that in addition to using the data rearrangement function to rearrange each compressed data matrix block, each index matrix block, each metadata matrix block, and each second correlation matrix block in shared memory, it is also possible to adopt the method of adding data padding to each compressed data matrix block, each index matrix block, each metadata matrix block, and each second correlation matrix block in shared memory to avoid bank conflicts.
[0124] S460. Perform secondary partitioning on each compressed data matrix block, each index matrix block, and each metadata matrix block in shared memory.
[0125] In the embodiments of the present invention, in order to adapt to the storage limitations of hardware registers, secondary partitioning can be performed on each compressed data matrix block, each index matrix block, and each metadata matrix block in shared memory.
[0126] Furthermore, when performing secondary partitioning on each compressed data matrix block, each index matrix block, and each metadata matrix block in shared memory, the secondary partitioning of the compressed data matrix block can also be performed first. During the process of secondary partitioning of the compressed data matrix block, in order to further consider the efficiency of subsequent data transfer and multiplication calculations, it can be set that during the process of secondary partitioning of the compressed data matrix block in shared memory, V is also an integer multiple of the segmentation size in the column segmentation direction. For example, secondary segmentation can be selected in the column segmentation direction or no secondary segmentation can be selected (only segmentation is performed in the row direction).
[0127] For example, assume that in global memory, when the compressed data matrix block is secondarily segmented into multiple new data blocks of size m b1 *k b1 it is required that V is an integer multiple of k b1 For example, k b1 is V / 4 or V / 8, etc.
[0128] Similarly, since there is a data or position correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix, after the secondary partitioning of the compressed data matrix block is completed, the secondary partitioning of the index matrix block and the metadata matrix block can be performed accordingly.
[0129] S470. Perform secondary partitioning on each second correlation matrix block in shared memory.
[0130] S480. According to the register mapping relationship matching the sparse computing unit, at least one of the secondary blockings of each compressed data matrix, the secondary blockings of each index matrix, and the secondary blockings of each metadata matrix is successively transferred from the shared memory to the hardware registers.
[0131] For a more intuitive understanding, in Figure 6 a block diagram showing an implementation of a data blocking and a multi-level transfer process applicable to the embodiments of the present invention is shown. In a specific example, taking the second sparse matrix as an example, as Figure 6 shown, it describes that after the compressed data matrix A in the first sparse matrix and the second sparse matrix B are respectively blocked in the global memory and the shared memory, and are successively transferred in the order of global memory -> shared memory -> register, after the calculation is completed by the sparse computing unit and gradually written back, the corresponding multiplication result matrix C is obtained in the global memory.
[0132] In this example, an example of data blocking and data transfer of a matrix multiplication of any size C(M×N)=A(M×K)×B(K×N) is given. In the embodiments of the present invention, in order to implement matrix multiplication calculation, the second sparse matrix B is row-column interchanged, that is, an array is selected to describe the positions where the effective data columns are located in the second sparse matrix B.
[0133] Specifically, in the global memory, since only the positions where the effective data columns (rows) are located are included in the selection array, the total amount of data in the selection array is len d pieces, and further the number of n-dimensional valid values in the second sparse matrix B is len d pieces. The compressed data matrix A is evenly divided into several pieces in the m dimension with m b as the size. The second sparse matrix B is divided in the n dimension, and the number of effective column vectors in each block data is the same as the data block divided with the size of n b in the selection array. Considering the storage size limitation of the computing device, the compressed data matrix A and the second sparse matrix B will be further evenly divided in the k direction with k b as the size. The multiplication results of multiple sub-blocks in the k direction are accumulated to obtain the final multiplication result matrix C block result. Based on this data division strategy, the size of the multiplication result matrix C calculated for each block data is m b ×n b , and the result calculated by the block calculation is a part of the final result of the multiplication result matrix C, which corresponds to the same row offset of the data block of the compressed data matrix A in the compressed data matrix A and the position where the data block of the second sparse matrix B has the same column offset in the selection array. The parallelism of each sparse operator is improved by the method that each thread group is responsible for one block.
[0134] Similarly, in shared memory, it also involves the process of re - chunking each data chunk transferred from global memory, moving it to registers for storage, and then performing matrix multiplication calculations by an adapted sparse computing unit. This process will not be elaborated here. In this embodiment, the matrix specification of the compressed data matrix A processed by the sparse computing unit once is mi * ki, and the matrix specification of the second sparse matrix B is ki * ni. For more convenient subsequent calculations, the aforementioned V is an integer multiple of ki.
[0135] In an alternative embodiment of the present invention, the re - chunking of each compressed data matrix, the re - chunking of each index matrix, and the re - chunking of each metadata matrix can all be moved to the register for calculation. Or, considering that register resources are very precious, only the re - chunking of each compressed data matrix and the re - chunking of each metadata matrix can be moved to the register for calculation, while the index matrix is retained in shared memory and only fetched and used from shared memory when needed.
[0136] It should be emphasized that, different from the calculation logic of general hardware computing units, when using a sparse computing unit to accelerate multiplication calculations, the data required for calculation needs to be placed in an adapted instruction register according to the requirements of the sparse instruction set. Therefore, in each embodiment of the present invention, the data in shared memory needs to be moved to a matching hardware register according to the register mapping relationship matching the sparse computing unit.
[0137] Optionally, the register mapping relationship can be read from the specification file configured at the factory of the computing chip. The register hardware file describes the mapping relationship between the calculation data at different positions and the hardware registers.
[0138] Specifically, Figure 7 shows a schematic diagram of a register mapping relationship applicable to the embodiments of the present invention, and this register mapping relationship is adapted to the re - chunking of the compressed data matrix. In a specific example, as Figure 7 shown, the T0{a0, a1} at the upper - left corner position in this register mapping relationship represents: the data in the 0th and 1st columns of the 0th row in the re - chunking of the compressed data matrix are allocated to the register a0 and register a1 matching the thread T0. Similarly, the T 0…3 {a4, a5} at the last position in the 0th row represents: the data in the 8th - 15th columns of the 0th row in the re - chunking of the compressed data matrix are respectively allocated to the register a4 and register a5 matching the thread T0, thread T1, thread T2, and thread T3. Among them, different threads T are pre - assigned to different sparse computing units for use.
[0139] Similarly, for the second block division of the metadata matrix and the second block division of the index matrix, the sparse computing unit also has a matching register mapping relationship. Based on the above register mapping relationship, each calculation data can be sequentially transferred from the shared memory to the hardware register.
[0140] In an optional implementation manner of the embodiment of the present invention, according to the register mapping relationship matching the sparse computing unit, a preset hardware instruction can be called to sequentially transfer each compressed data matrix second block division and each metadata matrix second block division, or each compressed data matrix second block division, each index matrix second block division, and each metadata matrix second block division from the shared memory to the hardware register, so as to achieve hardware acceleration of data loading.
[0141] Among them, the hardware instruction is associated with the specifications of the computing chip and can be queried and obtained from the instruction library of the computing chip. For example, it can be a 1dmatrix instruction, etc.
[0142] S490. According to the register mapping relationship matching the sparse computing unit, each second associated matrix second block division is sequentially loaded from the shared memory to the hardware register.
[0143] As shown before, by querying and obtaining the register mapping relationship set by the sparse computing unit for the second sparse matrix second block division, each second sparse matrix second block division can be sequentially loaded from the shared memory to the hardware register.
[0144] Similarly, in an optional implementation manner of this embodiment, according to the register mapping relationship matching the sparse computing unit, a preset hardware instruction can be called to sequentially load each second sparse matrix second block division from the shared memory to the hardware register, so as to achieve hardware acceleration of data loading.
[0145] It should be noted that in addition to the first sparse operator, when the second sparse operator and the third sparse operator are sequentially transferred from the global memory to the hardware register of the computing chip in the form of data block division, during the transfer process, the data in the original global memory needs to be rearranged or deformed.
[0146] S4100. Through the sparse computing unit of the computing chip, according to the data loaded in batches in the hardware register, the multiplication result of the first sparse matrix and the second associated matrix is gradually calculated.
[0147] S4110. Process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the corresponding functional function of the original mixture of experts model to obtain a sparse operator.
[0148] S4120. Update the original operator of the original mixture-of-experts model according to the sparse operator to obtain a target mixture-of-experts model.
[0149] Optionally, the multiplication result of the first sparse matrix and the second correlation matrix may be processed according to the processing logic of the broadcast multiplication function of the original mixture-of-experts model to obtain a first sparse operator; the multiplication result of the first sparse matrix and the second correlation matrix may be processed according to the processing logic of the activation function of the original mixture-of-experts model to obtain the second sparse operator; the multiplication result of the first sparse matrix and the second correlation matrix may be processed according to the processing logic of the dot product function of the original mixture-of-experts model to obtain a third sparse operator.
[0150] In an optional implementation manner of the embodiment of the present invention, through the sparse computing unit of the computing chip, according to the data loaded in batches in the hardware register, gradually calculating the multiplication result of the first sparse matrix and the second sparse matrix may include:
[0151] Through the sparse computing unit, obtain the current compressed data matrix quadratic block, the current metadata matrix quadratic block, and the current second correlation matrix quadratic block currently loaded in the hardware register;
[0152] Through the sparse computing unit, generate a first sparse block matrix according to the current metadata matrix quadratic block and the current compressed data matrix quadratic block;
[0153] Through the sparse computing unit, perform multiplication calculation according to the first sparse block matrix and the current second correlation matrix quadratic block.
[0154] As described above, when the first sparse matrix is compressed and stored, the compressed data matrix only stores the non-0 data in each valid row. In the actual first sparse matrix, each valid row contains at least one structured sparse computing storage unit in a set sparse mode. To ensure the accuracy of the multiplication calculation, it is necessary to first restore the current compressed data matrix quadratic block to the form of the aforementioned structured sparse computing storage unit. Since the metadata matrix stores the positions of the non-0 data in the structured sparse computing storage unit, furthermore, a first sparse block matrix can be generated according to the current metadata matrix quadratic block and the current compressed data matrix quadratic block.
[0155] That is to say, the first sparse block matrix can be understood as the matrix after restoring the structured sparse computing storage unit in the current compressed data matrix quadratic block. In a specific example, if the current compressed data matrix quadratic block is {A, B}, and the corresponding current metadata matrix quadratic block is {0, 2}, then a first sparse block matrix in the form of {A, 0, B, 0} can be restored.
[0156] Among them, through the sparse computing unit, according to a sparse block matrix and a current second correlation matrix for secondary blocking, a multiplication calculation is performed, which may specifically include: through the sparse computing unit, storing the intermediate calculation amounts generated during the multiplication calculation into the intermediate registers of the computing chip; and through the sparse computing unit, performing a matching multiplication calculation by obtaining the intermediate calculation amounts from the intermediate registers.
[0157] In the prior art, when using a general computing unit to perform a multiplication calculation, the results of the matrix multiplication calculation can be directly accumulated, but in essence, it increases the calculation of a lot of redundant data. In contrast, in the sparse computing units of the embodiments of the present invention, when performing a sparse multiplication calculation, when the multiplication calculation iterates along the K direction, since the sparse data comes from different rows, to ensure correctness, when different sub-blocks move along the K direction, the output results need to be mapped to different rows. Traditionally, passing the output results to specific registers according to indexes may cause the output C matrix to be reloaded into the local memory, which has a significant impact on the performance of the operator.
[0158] To avoid the above problems, the embodiments of the present invention reduce the frequent memory transfer between the global memory and the registers by introducing an additional intermediate register C IR in this way. Among them, in Figure 8 shows a schematic structural diagram of a sparse computing unit cooperating with an intermediate register to implement a calculation applicable to the embodiments of the present invention. As Figure 8 shown, when the sparse computing unit performs a sparse matrix multiplication calculation on the current compressed data matrix secondary block ( Figure 8 A in), the current metadata matrix secondary block ( Figure 8 metadata in), and the previous second correlation matrix secondary block ( Figure 8 B in), for multiple result data, the required intermediate results are loaded into C IR during the calculation process, and the intermediate results are used for calculation, and after the calculation is completed, the intermediate results are stored back to the corresponding result data position C.
[0159] By adding a new intermediate register in the computing chip, which is a simple hardware improvement, the calculation performance can be greatly improved and the calculation efficiency can be improved during the process of the sparse computing unit performing a sparse matrix multiplication calculation.
[0160] Based on the above embodiments, if the second correlation matrix includes a second sparse matrix, after performing multiplication calculation by the sparse calculation unit according to the first sparse block matrix and the current second sparse matrix sub-blocking, it may further include: after determining that the intermediate register stores the complete multiplication calculation result, converting the complete multiplication calculation result into a second sparse block matrix according to the current index matrix sub-block matching the current compressed data matrix sub-blocking; writing the second sparse block matrix back to the global memory of the computing chip level by level.
[0161] If the second correlation matrix includes a dense matrix corresponding to the second sparse matrix, after performing multiplication calculation by the sparse calculation unit according to the first sparse block matrix and the current second correlation matrix sub-blocking, it may further include: after determining that the intermediate register stores the complete multiplication calculation result, writing the multiplication calculation result back to the global memory of the computing chip level by level in the original sparse format.
[0162] Furthermore, since the sparse rows in the first sparse matrix are not considered when calculating the complete multiplication calculation result, therefore, if the right operand uses the second sparse matrix, after obtaining the complete multiplication calculation result, it is necessary to combine the current index matrix sub-block matching the current compressed data matrix sub-blocking, and after adding the matching sparse rows to the complete multiplication calculation result, obtain the second sparse block matrix as the real block calculation result. As described above, the current index matrix sub-block can also be transferred to the hardware register or only stored in the shared memory, and this embodiment does not limit this. If the right operand uses the dense matrix corresponding to the second sparse matrix, after obtaining the complete multiplication calculation result, the multiplication calculation result can be directly written back to the global memory of the computing chip level by level in the original sparse format without further sparse processing of the multiplication calculation result. As described above, the multiplication calculation result can be transferred to the hardware register in the original sparse format or only stored in the shared memory, and this embodiment does not limit this.
[0163] In summary, in combination with the model architecture, matrix multiplication operations with sparse attributes are selected from the calculation flow of the mixture-of-experts model. Sparse comes from two aspects. On the one hand, it comes from the sparsity of the weights. The redundant attributes in its parameters enable the calculation process to skip the calculation of less important parameters. On the other hand, sparsity can also come from the mixture-of-experts model. The input values are selected through a gating circuit to calculate the matching expert parameters. From the perspective of the expert, the input values have sparse characteristics, and only some of the inputs participate in the calculation of the current expert. The selection of such sparse operators effectively improves the execution speed of the mixture-of-experts model without affecting the model accuracy. Based on the strategy of operator fusion in the sparse data structure and data stream execution, referring to the compilation optimization scheme of sparse operators, specific sparse operators in the sparse operator library are selected and injected into the artificial intelligence framework and compiler to replace the dense matrix multiplication operation of the original mixture-of-experts model. For example, the two matrix multiplication operations that accept the original input in the expert architecture of the mixture-of-experts model will be injected with operators in the double-sparse mode, while the subsequent matrix multiplication that accepts the intermediate activation values will be injected with unidirectional sparse (sparse × dense) operators.
[0164] The technical solution of the embodiment of the present invention fully considers the sparse characteristics when performing matrix multiplication calculation based on a sparse matrix. By adopting data partitioning and data transfer techniques adapted to the sparse characteristics, the operator optimization of the mixture-of-experts model is realized, and the hardware acceleration performance of the sparse computing unit in the computing chip is fully utilized in the model calculation scenario of the mixture-of-experts model, thereby greatly optimizing the overhead of computing, bandwidth, and storage resources in the sparse matrix multiplication operation process.
[0165] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analysis data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards in the relevant regions.
[0166] It should be noted that any permutation and combination of the technical features in the above embodiments also fall within the protection scope of the present invention.
[0167] Figure 9 is a schematic diagram of an operator optimization device for a mixture-of-experts model provided by an embodiment of the present invention, as Figure 9 shown, the device includes: a sparse operator generation module 910 and a sparse operator update module 920, wherein:
[0168] The sparse operator generation module 910 is used to generate sparse operators applicable to the original mixture-of-experts model by combining the sparse computing unit of the computing chip in a data partitioning manner; wherein, the number of the sparse operators is multiple;
[0169] The sparse operator update module 920 is configured to update the original operator of the original mixture-of-experts model according to the sparse operator, so as to obtain a target mixture-of-experts model.
[0170] In the embodiment of the present invention, by adopting a data chunking method in combination with the sparse computing unit of the computing chip, a plurality of sparse operators suitable for the original mixture-of-experts model are generated, so as to update the original operator of the original mixture-of-experts model according to the generated sparse operators, and a target mixture-of-experts model is obtained. The sparse characteristics during matrix multiplication calculation of the mixture-of-experts model based on the sparse matrix are fully considered. By adopting data chunking and data transfer techniques adapted to the sparse characteristics, the optimization of the operators of the mixture-of-experts model is realized, and the hardware acceleration performance of the sparse computing unit in the computing chip is fully exerted in the model calculation scenario of the mixture-of-experts model, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the sparse matrix multiplication operation process.
[0171] Optionally, the sparse operator includes a first sparse operator, and the sparse operator generation module 910 is further configured to: locate a set of compressed matrices and a second associated matrix matching the first sparse matrix in the global memory of the computing chip according to the first sparse matrix multiplication requirement; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; wherein, the first sparse matrix includes a plurality of compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of the valid rows in the first sparse matrix in the corresponding compression units, and a metadata matrix for storing the positions of the non-zero data in the corresponding structured sparse computing storage units; transfer the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data chunking form step by step; through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the broadcast multiplication function of the original mixture-of-experts model to obtain the first sparse operator.
[0172] Optionally, the sparse operator includes a second sparse operator, and the sparse operator generation module 910 is further configured to: locate a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip according to the second sparse matrix multiplication requirement; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; wherein, the first sparse matrix includes a plurality of compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in their respective compression units, and a metadata matrix for storing the positions of non-zero data in their respective structured sparse computing storage units; transfer the set of compressed matrices and the second associated matrix from the global memory to the hardware registers of the computing chip in a data-block form, and during the transfer, rearrange or transform the data in the original global memory; through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the activation function of the original mixture-of-experts model to obtain the second sparse operator.
[0173] Optionally, the sparse operator includes a third sparse operator, and the sparse operator generation module 910 is further configured to: locate a set of compressed matrices and a second associated matrix that match the first sparse matrix in the global memory of the computing chip according to the third sparse matrix multiplication requirement; wherein, the second associated matrix includes a second sparse matrix or a dense matrix corresponding to the second sparse matrix; wherein, the first sparse matrix includes a plurality of compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of the valid rows in the first sparse matrix in their respective compression units, and a metadata matrix for storing the positions of the non-zero data in their respective structured sparse computing storage units; transfer the set of compressed matrices and the second associated matrix to the hardware registers of the computing chip in a data-block form level by level from the global memory, and rearrange or transform the data in the original global memory during the transfer process; through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second associated matrix according to the data loaded in batches in the hardware registers; process the multiplication result of the first sparse matrix and the second associated matrix according to the processing logic of the dot product function of the original mixture-of-experts model to obtain the third sparse operator.
[0174] Optionally, the sparse operator generation module 910 is further configured to: divide the compressed data matrix in the global memory into data blocks, and match the index matrix and the metadata matrix with data blocks according to the data correspondence between the compressed data matrix, the index matrix, and the metadata matrix; sequentially load each compressed data matrix block, each index matrix block, and each metadata matrix block from the global memory to the shared memory; perform secondary block division on each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory; according to the register mapping relationship matching the sparse computing unit, sequentially transfer at least one of each compressed data matrix secondary block, each index matrix secondary block, and each metadata matrix secondary block from the shared memory to the hardware registers.
[0175] Optionally, the sparse operator generation module 910 is further configured to: locate a selection array matching the second correlation matrix in the computing chip, where the selection array is used to describe the positions of valid rows in the second correlation matrix; the selection array of the second sparse matrix is the index of the non-zero data positions therein, and the selection array of the dense matrix corresponding to the second sparse matrix is the position mapping of the data in the dense matrix in the second sparse matrix; according to the selection array, perform data chunking on the second correlation matrix in the global memory; for each chunk of the second correlation matrix, sequentially transfer it from the global memory to the shared memory; perform secondary chunking on each chunk of the second correlation matrix in the shared memory; according to the register mapping relationship matching the sparse computing unit, for each secondary chunk of the second correlation matrix, sequentially load it from the shared memory to the hardware register.
[0176] Optionally, the sparse operator generation module 910 is further configured to: use a preset data rearrangement function to rearrange each chunk of the compressed data matrix, each chunk of the index matrix, and each chunk of the metadata matrix in the shared memory to avoid bank conflicts; use the preset data rearrangement function to rearrange each chunk of the second correlation matrix in the shared memory to avoid bank conflicts.
[0177] Optionally, the sparse operator generation module 910 is further configured to: according to the register mapping relationship matching the sparse computing unit, call a preset hardware instruction to perform secondary chunking on each chunk of the compressed data matrix and each chunk of the metadata matrix, or perform secondary chunking on each chunk of the compressed data matrix, each chunk of the index matrix, and each chunk of the metadata matrix, and sequentially transfer them from the shared memory to the hardware register to achieve hardware acceleration of data loading; according to the register mapping relationship matching the sparse computing unit, call a preset hardware instruction to perform secondary chunking on each chunk of the second correlation matrix and sequentially load it from the shared memory to the hardware register to achieve hardware acceleration of data loading.
[0178] Optionally, the sparse operator generation module 910 is further configured to: through the sparse computing unit, obtain the current secondary chunk of the compressed data matrix, the current secondary chunk of the metadata matrix, and the current secondary chunk of the second correlation matrix currently loaded in the hardware register; through the sparse computing unit, generate a primary sparse chunk matrix according to the current secondary chunk of the metadata matrix and the current secondary chunk of the compressed data matrix; through the sparse computing unit, perform multiplication calculation according to the primary sparse chunk matrix and the current secondary chunk of the second correlation matrix.
[0179] Optionally, the sparse operator generation module 910 is further configured to: store the intermediate computation quantities generated during the multiplication computation into the intermediate registers of the computing chip through the sparse computing unit; and perform the matching multiplication computation by obtaining the intermediate computation quantities from the intermediate registers through the sparse computing unit.
[0180] Optionally, the sparse operator generation module 910 is further configured to: if the second association matrix includes a second sparse matrix, after determining that the complete multiplication computation result is stored in the intermediate register, convert the complete multiplication computation result into a quadratic sparse block matrix according to the current index matrix quadratic block that matches the quadratic block division of the current compressed data matrix; write the quadratic sparse block matrix back to the global memory of the computing chip level by level; if the second association matrix includes the dense matrix corresponding to the second sparse matrix, after determining that the complete multiplication computation result is stored in the intermediate register, write the multiplication computation result back to the global memory of the computing chip level by level in the original sparse format.
[0181] Optionally, the first sparse matrix specifically includes a plurality of compression units of size M*V; each of the compression units specifically includes d valid rows, where 1≤d<M; and each of the structured sparse computing and storage units is in an N:L sparse mode, where L is the total number of data included in the structured sparse computing and storage unit, and N is the number of non-zero data included in the structured sparse computing and storage unit.
[0182] Optionally, during each data block division process, V is an integer multiple of the division size in the column division direction; and for each sparse operator, the matrix specification of the first sparse matrix processed once is mi*ki, where V is an integer multiple of ki.
[0183] Optionally, the first sparse matrix is the sparsified model weight matrix in the target mixture-of-experts model, the second sparse matrix is the intermediate activation value sparse matrix input to the mixture model layer in the target mixture-of-experts model, and the dense matrix corresponding to the second sparse matrix is the dense matrix of the intermediate activation value input to the mixture model layer in the target mixture-of-experts model.
[0184] The operator optimization device of the above mixture-of-experts model can execute the operator optimization method of the mixture-of-experts model provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the operator optimization method of the mixture-of-experts model provided in any embodiment of the present invention.
[0185] Since the operator optimization device of the mixture-of-experts model introduced above is a device that can execute the operator optimization method of the mixture-of-experts model in the embodiments of the present invention, based on the operator optimization method of the mixture-of-experts model introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the operator optimization device of the mixture-of-experts model in this embodiment. Therefore, the implementation of how the operator optimization device of the mixture-of-experts model realizes the operator optimization method of the mixture-of-experts model in the embodiments of the present invention will not be described in detail herein. As long as the device adopted by those skilled in the art to implement the operator optimization method of the mixture-of-experts model in the embodiments of the present invention falls within the scope of protection of this application.
[0186] Figure 10 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0187] As Figure 10 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0188] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0189] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the operator optimization method of the mixture of experts model.
[0190] Optionally, the operator optimization method of the mixture of experts model may include: combining the sparse computing units of the computing chip in a data-blocking manner to generate sparse operators applicable to the original mixture of experts model; wherein the number of the sparse operators is multiple; and updating the original operators of the original mixture of experts model according to the sparse operators to obtain a target mixture of experts model.
[0191] In some embodiments, the operator optimization method of the mixture of experts model can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the operator optimization method of the mixture of experts model described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the operator optimization method of the mixture of experts model in any other suitable manner (e.g., by means of firmware).
[0192] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special-purpose or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0194] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0195] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0196] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0197] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0198] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0199] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An operator optimization method of a hybrid expert model, characterized in that: include: A sparse operator suitable for the original hybrid expert model is generated by combining data block with a sparse computing unit of a computing chip; wherein the number of the sparse operators is multiple; updating the original operator of the original hybrid expert model according to the sparse operator to obtain a target hybrid expert model; Among them, the data block method is combined with the sparse computing unit of the computing chip to generate a sparse operator suitable for the original hybrid expert model, including: According to the sparse matrix multiplication requirement, a set of compressed matrices matching the first sparse matrix and a second association matrix are located in the global memory of the computing chip; wherein the second association matrix includes the second sparse matrix or a dense matrix corresponding to the second sparse matrix; Wherein, the first sparse matrix comprises a plurality of compression units of set size, each of the compression units comprises at least one valid row and at least one sparse row, each of the valid row comprises at least one structured sparse computing storage unit; each of the structured sparse computing storage units is in the same sparse mode; the compression matrix set comprises a compression data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the position of the valid rows in the first sparse matrix in the corresponding compression unit, and a metadata matrix for storing the position of the non-zero data in the corresponding structured sparse computing storage unit; the sparse rows are the matrix rows in which all the data are 0, and the sparse mode is the proportion of non-zero value data in the structured sparse computing storage unit to all the data; The compressed matrix set and the second association matrix are transferred from the global memory to the hardware register of the computing chip in the form of data blocks step by step; The sparse computing unit of the computing chip gradually calculates the multiplication result of the first sparse matrix and the second association matrix according to the data loaded in batches in the hardware register; The multiplication result of the first sparse matrix and the second correlation matrix is processed according to the processing logic of the corresponding functional function of the original hybrid expert model to obtain the sparse operator.
2. The method according to claim 1, characterized in that The sparse operator includes a first sparse operator, and the sparse computing unit of the computing chip is combined with the data block method to generate a sparse operator suitable for the original hybrid expert model, including: According to the first sparse matrix multiplication requirement, locating a set of compressed matrices and a second associative matrix matching the first sparse matrix in a global memory of a computing chip; The compressed matrix set and the second association matrix are transferred from the global memory to the hardware register of the computing chip in the form of data blocks step by step; The sparse computing unit of the computing chip gradually calculates the multiplication result of the first sparse matrix and the second association matrix according to the data loaded in batches in the hardware register; The multiplication result of the first sparse matrix and the second association matrix is processed according to the processing logic of the broadcast multiplication function of the original hybrid expert model to obtain the first sparse operator.
3. The method according to claim 1, characterized in that The sparse operator includes a second sparse operator, and the sparse computing unit of the computing chip is combined with the data block method to generate a sparse operator suitable for the original hybrid expert model, including: According to the second sparse matrix multiplication requirement, locating a set of compressed matrices and a second association matrix matching the first sparse matrix in the global memory of the computing chip; The compressed matrix set and the second association matrix are transferred from the global memory to the hardware register of the computing chip in the form of data blocks, and during the transfer process, the data in the original global memory is rearranged or deformed; The sparse computing unit of the computing chip gradually calculates the multiplication result of the first sparse matrix and the second association matrix according to the data loaded in batches in the hardware register; The multiplication result of the first sparse matrix and the second association matrix is processed according to the processing logic of the activation function of the original hybrid expert model to obtain the second sparse operator.
4. The method according to claim 1, characterized in that: The sparse operator includes a third sparse operator, and the sparse computing unit of the computing chip is combined with the data block method to generate a sparse operator suitable for the original hybrid expert model, including: According to the third sparse matrix multiplication requirement, locating a set of compressed matrices and a second association matrix matching the first sparse matrix in the global memory of the computing chip; The compressed matrix set and the second association matrix are transferred from the global memory to the hardware register of the computing chip in the form of data blocks, and during the transfer process, the data in the original global memory is rearranged or deformed; The sparse computing unit of the computing chip gradually calculates the multiplication result of the first sparse matrix and the second association matrix according to the data loaded in batches in the hardware register; The multiplication result of the first sparse matrix and the second correlation matrix is processed according to the processing logic of the point multiplication function of the original hybrid expert model to obtain the third sparse operator.
5. The method according to any one of claims 2 to 4, characterized in that: The compressed matrix set is transferred from the global memory to the hardware register of the computing chip in the form of data blocks, including: The compressed data matrix in the global memory is divided into data blocks, and the index matrix and the metadata matrix are matched into data blocks according to the data correspondence between the compressed data matrix, the index matrix and the metadata matrix; Loading each compressed data matrix block, each index matrix block, and each metadata matrix block from the global memory to the shared memory one by one; Secondarily partitioning each of the compressed data matrix blocks, each of the index matrix blocks, and each of the metadata matrix blocks in the shared memory; According to the register mapping relationship matching the sparse computing unit, at least one item of each of the compressed data matrix secondary blocks, each of the index matrix secondary blocks and each of the metadata matrix secondary blocks is moved from the shared memory to the hardware register one by one.
6. The method according to claim 5, characterized in that While locating the set of compressed matrices matching the first sparse matrix and the second association matrix in the global memory of the computing chip, the method further includes: Locating a selection array matching the second incidence matrix in the computing chip, wherein the selection array is used to describe valid row positions in the second incidence matrix; the selection array of the second sparse matrix is the index of the non-zero data position therein, and the selection array of the dense matrix corresponding to the second sparse matrix is the position mapping of the data in the dense matrix in the second sparse matrix; Accordingly, the second association matrix is moved from the global memory to the hardware register of the computing chip in the form of data blocks step by step, including: According to the selection array, dividing the second association matrix in the global memory into data blocks; Divide each second association matrix into blocks, and transfer them from the global memory to the shared memory one by one; Perform secondary block division on each of the second association matrix blocks in the shared memory; According to the register mapping relationship matching the sparse computing unit, each second association matrix is divided into two blocks and loaded from the shared memory to the hardware register one by one.
7. The method according to claim 6, characterized in that After loading each compressed data matrix block, each index matrix block, and each metadata matrix block from the global memory to the shared memory one by one, the method further includes: Using a preset data rearrangement function, rearrange the compressed data matrix blocks, the index matrix blocks, and the metadata matrix blocks in the shared memory to avoid storage bank conflicts; and After dividing each second correlation matrix into blocks and transferring them from the global memory to the shared memory one by one, the method further includes: The preset data rearrangement function is used to rearrange each of the second association matrix blocks in the shared memory to avoid storage bank conflicts.
8. The method according to claim 6, characterized in that According to the register mapping relationship matched with the sparse computing unit, at least one item of each of the compressed data matrix secondary blocks, each of the index matrix secondary blocks, and each of the metadata matrix secondary blocks is moved from the shared memory to the hardware register one by one, including: According to the register mapping relationship matched with the sparse computing unit, calling the preset hardware instruction, each of the compressed data matrix secondary blocks and each of the metadata matrix secondary blocks, or each of the compressed data matrix secondary blocks, each of the index matrix secondary blocks and each of the metadata matrix secondary blocks, are moved from the shared memory to the hardware register one by one, so as to realize the hardware acceleration of data loading; and According to the register mapping relationship matching the sparse computing unit, each second association matrix is divided into blocks twice, and loaded from the shared memory to the hardware register one by one, including: According to the register mapping relationship matching the sparse computing unit, the preset hardware instructions are called to divide each of the second association matrices into blocks twice, and load them from the shared memory to the hardware registers one by one to achieve hardware acceleration of data loading.
9. The method according to claim 6, characterized in that The multiplication result of the first sparse matrix and the second association matrix is calculated step by step according to the data loaded in batches in the hardware register by the sparse computing unit of the computing chip, including: Obtaining, by means of the sparse computing unit, the current compressed data matrix secondary block, the current metadata matrix secondary block, and the current second association matrix secondary block currently loaded in the hardware register; Generate a primary sparse block matrix according to the secondary block of the current metadata matrix and the secondary block of the current compressed data matrix by the sparse computing unit; The sparse computing unit performs multiplication calculation according to the primary sparse block matrix and the secondary block of the current second correlation matrix.
10. The method according to claim 9, characterized in that The sparse computing unit performs a multiplication calculation according to the primary sparse block matrix and the secondary block of the current second correlation matrix, specifically including: The intermediate calculation amount generated in the multiplication calculation process is stored in the intermediate register of the computing chip through the sparse computing unit; and The sparse computing unit performs a matching multiplication calculation by acquiring the intermediate calculation amount from the intermediate register.
11. The method according to claim 10, characterized in that If the second incidence matrix includes a second sparse matrix, after performing multiplication calculation according to the first sparse block matrix and the second current incidence matrix second block by the sparse computing unit, the method further includes: After determining that the intermediate register stores a complete multiplication calculation result, converting the complete multiplication calculation result into a secondary sparse block matrix according to a current index matrix secondary block matching the current compressed data matrix secondary block; Writing the secondary sparse block matrix back to the global memory of the computing chip step by step; If the second incidence matrix includes a dense matrix corresponding to the second sparse matrix, after performing multiplication calculation according to the first sparse block matrix and the second second incidence matrix second block by the sparse computing unit, the method further includes: After determining that the complete multiplication calculation result is stored in the intermediate register, the multiplication calculation result is written back to the global memory of the computing chip step by step according to the original sparse format.
12. The method according to any one of claims 2 to 4, characterized in that: The first sparse matrix specifically includes a plurality of compression units of size M*V; Specifically, each compression unit includes d valid rows, wherein 1≤d<M; and Each of the structured sparse computing storage units is in a sparse mode of N:L, where L is the total amount of data contained in the structured sparse computing storage unit, and N is the amount of non-zero data contained in the structured sparse computing storage unit.
13. The method according to claim 12, characterized in that In each data segmentation process, V is an integer multiple of the segmentation size in the column segmentation direction; and The matrix specification of the first sparse matrix processed once in each of the sparse operators is mi*ki, where V is an integer multiple of ki.
14. The method according to any one of claims 2 to 4, characterized in that: The first sparse matrix is the sparse model weight matrix in the target hybrid expert model, the second sparse matrix is the sparse matrix of intermediate activation values input to the hybrid model layer in the target hybrid expert model, and the dense matrix corresponding to the second sparse matrix is the dense matrix of intermediate activation values input to the hybrid model layer in the target hybrid expert model.
15. An operator optimization device for a hybrid expert model, characterized in that: include: A sparse operator generation module, used to generate a sparse operator suitable for the original hybrid expert model by combining the sparse computing unit of the computing chip in a data block manner; wherein the number of the sparse operators is multiple; A sparse operator updating module, used for updating the original operator of the original hybrid expert model according to the sparse operator to obtain a target hybrid expert model; Wherein, the sparse operator generation module is also used for: According to the sparse matrix multiplication requirement, a set of compressed matrices matching the first sparse matrix and a second association matrix are located in the global memory of the computing chip; wherein the second association matrix includes the second sparse matrix or a dense matrix corresponding to the second sparse matrix; Wherein, the first sparse matrix comprises a plurality of compression units of set size, each of the compression units comprises at least one valid row and at least one sparse row, each of the valid row comprises at least one structured sparse computing storage unit; each of the structured sparse computing storage units is in the same sparse mode; the compression matrix set comprises a compression data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the position of the valid rows in the first sparse matrix in the corresponding compression unit, and a metadata matrix for storing the position of the non-zero data in the corresponding structured sparse computing storage unit; the sparse rows are the matrix rows in which all the data are 0, and the sparse mode is the proportion of non-zero value data in the structured sparse computing storage unit to all the data; The compressed matrix set and the second association matrix are transferred from the global memory to the hardware register of the computing chip in the form of data blocks step by step; The sparse computing unit of the computing chip gradually calculates the multiplication result of the first sparse matrix and the second association matrix according to the data loaded in batches in the hardware register; The multiplication result of the first sparse matrix and the second correlation matrix is processed according to the processing logic of the corresponding functional function of the original hybrid expert model to obtain the sparse operator.
16. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the operator optimization method of the hybrid expert model described in any one of claims 1-14.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the operator optimization method of the hybrid expert model described in any one of claims 1 to 14 when executed.
18. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the operator optimization method of the hybrid expert model described in any one of claims 1 to 14 is implemented.