Storage Optimization Method, Device, Equipment, Medium and Product of Mixture-of-Experts Model
By converting the sparse model weight matrix of the hybrid expert model into compressed data, index and metadata matrix forms, the problem of inefficient computing efficiency of the hybrid expert model is solved, and the optimization of storage resources and the improvement of computing efficiency is achieved.
Patent Information
- Application Number
- CN202510308467.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing sparse compression and orchestration optimization techniques fail to effectively optimize double sparse features during the calculation process of hybrid expert models, resulting in ineffective computing efficiency and waste of storage resources.
The sparse model weight matrix of the hybrid expert model is converted into the form of compressed data matrix, index matrix and metadata matrix, and stored as a set of compressed matrixes, optimize storage resource utilization, and improve computing efficiency through sparse computing hardware.
It effectively reduces the storage resource overhead and bandwidth overhead of the computing chip, improves computing efficiency, and makes full use of the computing advantages of sparse computing hardware.
Smart Images

Figure CN119808861B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, medium and product for optimizing the storage of a mixture of experts model. Background Art
[0002] The sparse data representation and layout optimization technology is a technical solution that can effectively optimize the computational, bandwidth, and storage resource overheads during model training, inference, and fine-tuning.
[0003] In the prior art, when performing model calculations on a mixture of experts model, it mainly involves sparsely compressing the sparse weight data to be calculated and the input sparse intermediate activation data in a specified sparse representation format, and arranging them in a specified storage format in memory for subsequent calculations.
[0004] In the process of implementing the present invention, the inventors found that the calculation between the sparse intermediate activation data and the sparse weight data belongs to double-sparse calculation. However, the current mainstream sparse compression and layout optimization technologies lack an efficient representation of the double-sparse features in the calculation process of the mixture of experts model. Therefore, directly using the existing sparse compression and layout optimization schemes cannot fully optimize the calculation process of the mixture of experts model. Summary of the Invention
[0005] Embodiments of the present invention provide a method, device, equipment, medium and product for optimizing the storage of a mixture of experts model, and creatively propose an efficient data compression method for the sparsified model weight matrix in the mixture of experts model to effectively optimize the storage resource overhead of the mixture of experts model.
[0006] According to one aspect of the embodiments of the present invention, there is provided a method for optimizing the storage of a mixture of experts model, including:
[0007] Loading a target mixture of experts model into a computing chip, and obtaining each sparsified model weight matrix in the target mixture of experts model;
[0008] Wherein, the sparsified model weight matrix includes multiple compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode;
[0009] Generating a set of compression matrices corresponding to each model weight matrix respectively;
[0010] Among them, the set of compression matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix within their respective compression units, and a metadata matrix for storing the positions of non-zero data within their respective structured sparse computing storage units;
[0011] The sparsified model weight matrices in the target mixture-of-experts model are respectively stored as a matching set of compression matrices to achieve storage optimization of the target mixture-of-experts model.
[0012] According to another aspect of the embodiments of the present invention, there is also provided a storage optimization device for a mixture-of-experts model, including:
[0013] A model weight matrix acquisition module, configured to load the target mixture-of-experts model into a computing chip and acquire the sparsified model weight matrices in the target mixture-of-experts model;
[0014] Among them, the sparsified model weight matrix includes multiple compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode;
[0015] A compression matrix set generation module, configured to generate a set of compression matrices respectively corresponding to each model weight matrix;
[0016] Among them, the set of compression matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix within their respective compression units, and a metadata matrix for storing the positions of non-zero data within their respective structured sparse computing storage units;
[0017] A storage optimization module, configured to store the sparsified model weight matrices in the target mixture-of-experts model as matching sets of compression matrices respectively to achieve storage optimization of the target mixture-of-experts model.
[0018] According to another aspect of the embodiments of the present invention, there is provided an electronic device, where the electronic device includes:
[0019] At least one computing chip; and
[0020] A memory communicatively connected to the at least one computing chip; where
[0021] The memory stores a computer program executable by the at least one computing chip, and when the computer program is executed by the at least one computing chip, the at least one computing chip can execute the storage optimization method for the mixture-of-experts model according to any embodiment of the present invention.
[0022] On the other hand, according to an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computing chip, the storage optimization method of the mixture-of-experts model according to any embodiment of the present invention is implemented.
[0023] On the other hand, according to an embodiment of the present invention, a computer program product is further provided, including computer instructions. When the computer instructions are executed by a computing chip, the steps of the storage optimization method of the mixture-of-experts model according to any embodiment of the present invention are implemented.
[0024] In the technical solution of the embodiment of the present invention, by respectively converting the sparsified model weight matrices in specific formats in the target mixture-of-experts model into a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing storage unit, and then performing optimized storage in the computing chip for implementing the mixture-of-experts model calculation, the storage resource overhead of the computing chip can be effectively reduced. In addition, when using the model weight matrix in the above data compression form for model calculation, the sparse computing hardware in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 is a flowchart of a storage optimization method for a mixture-of-experts model provided according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of a comparison structure between a model weight matrix and a matching set of compression matrices applicable to an embodiment of the present invention;
[0029] Figure 3 is a flowchart of another storage optimization method for a mixture-of-experts model provided by an embodiment of the present invention;
[0030] Figure 4 It is a comparison schematic diagram of the data rearrangement result required and matched by a sparse computing unit applicable to the embodiment of the present invention;
[0031] Figure 5 It is a comparison structural schematic diagram of an intermediate activation value sparse matrix and a matched selection array applicable to the embodiment of the present invention;
[0032] Figure 6 It is a structural schematic diagram of a storage optimization device for a mixture-of-experts model provided according to the embodiment of the present invention;
[0033] Figure 7 It is a structural schematic diagram of an electronic device for implementing the storage optimization method of the mixture-of-experts model in the embodiment of the present invention;
[0034] Figure 8 It is a structural diagram of a computing chip applicable to the embodiment of the present invention. Detailed implementation manners
[0035] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0037] Figure 1The flowchart of a storage optimization method of a hybrid expert model provided by an embodiment of the present invention is applicable to the case where each sparse model weight matrix in the hybrid expert model is converted into a matching compressed matrix set and then optimized for storage in a computing chip used to implement hybrid expert model calculation. The method can be executed by a storage optimization device for a hybrid expert model, which can be implemented in the form of hardware and / or software and can generally be configured in an electronic device where the computing chip is located. Figure 1 As shown, the method includes:
[0038] S110, loading the target hybrid expert model into the computing chip, and obtaining the sparse model weight matrix of each item in the target hybrid expert model.
[0039] Optionally, the computing chip can be understood as an integrated circuit used to implement a set computing task (for example, Internet of Things control, high-performance computing or mobile computing, etc.). The computing chip can be a general computing chip, such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated computing chip, such as an NPU (Neural Processing Unit), which is not limited in this embodiment.
[0040] The target hybrid expert model refers to a specific hybrid expert model that is loaded into a computing chip to implement model reasoning tasks. The hybrid expert model is an efficient neural network architecture that improves overall performance by combining multiple specialized sub-models. The core idea of this model is to integrate different "expert" networks together, with each expert playing a role in his or her area of expertise. These experts can be small multilayer perceptrons (MLP) or more complex large language models (LLM). The model usually contains a gating network that is responsible for deciding which expert should be activated and involved in the calculation of the output when processing a specific input.
[0041] Specifically, the hybrid expert model can be applied in the fields of computer vision, natural language processing, medical treatment, or autonomous driving;
[0042] The computer vision field includes image classification, target detection or face recognition; the natural language processing field includes text classification, machine translation or speech recognition, etc.
[0043] For the convenience of implementing subsequent model calculations, in this embodiment, the model weight matrices of each model layer in the target mixture-of-experts model are all processed into sparsified model weight matrices with a specific structure.
[0044] Among them, the sparsified model weight matrix includes multiple compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse pattern.
[0045] Specifically, each model weight matrix in the target mixture-of-experts model includes an integer number (one or more) of compression units, that is, the size (number of rows * number of columns) of each model weight matrix can be divisible by the size of the compression unit. The compression unit can be understood as a small matrix of a specific size, and this small matrix has at least two matrix rows. At least one of the above at least two matrix rows is a valid row and at least one is a sparse row. A valid row can be understood as a matrix row that includes at least one valid data (non-zero), and a sparse row refers to a matrix row where all data are 0.
[0046] Furthermore, each valid row in the compression unit includes an integer number (one or more) of structured sparse computing storage units. Among them, all the structured sparse computing storage units included in the model weight matrix are in the same sparse pattern. Generally speaking, for the convenience of calculation, the size of the structured sparse computing storage unit can be adapted to the computing scale of the hardware computing unit in the computing chip.
[0047] The sparse pattern can be understood as the proportion of valid data (non-zero value data) in all data. For example, the sparse pattern can be 2:4 or 4:8, etc. That is, for a 2:4 sparse pattern, the structured sparse computing storage unit contains a total of 4 data, and 2 of them are non-zero data. The arrangement positions of the above two non-zero data in the structured sparse computing storage unit are not restricted.
[0048] In this embodiment, two methods can be used to make the model weight matrices of each model layer in the target mixture-of-experts model in the above clearly defined special structure: one is to strongly constrain the model weight matrices in each model layer during the training stage of the target mixture-of-experts model, so that each model weight matrix meets the requirements of the above special structure; the other is to strongly constrain the model weight matrices in each model layer during the model fine-tuning stage after model pre-training, so that each model weight matrix meets the requirements of the above special structure.
[0049] In this embodiment, the reason for restricting each model weight matrix in the target mixture-of-experts model to the above special structure is to improve the calculation efficiency and better adapt to the specific hardware calculation units in the calculation chip when performing calculations based on this target mixture-of-experts model. The determination method of the above special structure will not be elaborated here. That is, the model weight matrix is adapted to the calculation performance of the sparse calculation unit in the calculation chip (the calculation scale of the calculation components included, the number of registers included, and the number of cache buffers, etc.).
[0050] S120. Generate a set of compression matrices corresponding to each model weight matrix respectively.
[0051] As mentioned above, in each model weight matrix, there are a large number of 0-valued data. If the above model weight matrices are directly stored in the calculation chip according to the original size of the model weight matrix, it will cause a waste of a large number of storage units. In addition, directly performing subsequent model calculations based on the model weight matrix of the original size has low calculation efficiency and cannot efficiently use the hardware calculation units in the calculation chip.
[0052] In view of this, in this embodiment, for the model weight matrix with the above special structure, a novel and efficient data compression method is creatively proposed to solve the above technical problems.
[0053] In this embodiment, through a specific data compression method, each model weight matrix can be tightly stored in the form of a set of compression matrices. Specifically, in the set of compression matrices, three compression matrices corresponding to the model weight matrix are stored, namely the compressed data matrix, the index matrix, and the metadata matrix.
[0054] Among them, the compressed data matrix is used to store the compressed data of the non-0 data in the model weight matrix, the index matrix is used to store the positions of the valid rows in the model weight matrix in the corresponding compression unit, and the metadata matrix is used to store the positions of the non-0 data in the structured sparse calculation storage unit.
[0055] Obviously, through the above data compression method, the specific position of each non-0 data in the original model weight matrix can be determined by three small matrices. This data compression method can also effectively reduce the consumption of storage resources in the calculation chip and facilitate calculations based on this model weight matrix.
[0056] S130. Store each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compression matrices to achieve storage optimization of the target mixture-of-experts model.
[0057] Specifically, when it is necessary to store the sparsified model weight matrices in the target mixture-of-experts model in the storage space (typically, global memory) of a computing chip, each set of compressed matrices obtained after compression is used to replace each sparsified model weight matrix for storage, so as to achieve storage optimization.
[0058] In the technical solution of the embodiment of the present invention, by respectively converting each sparsified model weight matrix in a specific format in the target mixture-of-experts model into a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing storage unit, and then performing optimized storage in a computing chip for implementing the calculation of the mixture-of-experts model, the storage resource overhead of the computing chip can be effectively reduced. In addition, when using the model weight matrix in the above data compression form for model calculation, the sparse computing hardware in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0059] In an optional implementation manner of this embodiment, each sparsified model weight matrix specifically includes multiple compression units of size M*V; each compression unit specifically includes d valid rows, where 1≤d<M; and each structured sparse computing storage unit is in a sparse mode of N:L, where L is the total number of data included in the structured sparse computing storage unit, and N is the number of non-zero data included in the structured sparse computing storage unit.
[0060] Correspondingly, generating a set of compressed matrices corresponding to each model weight matrix may include:
[0061] Obtain the current model weight matrix of size m*k being processed, where m is an integer multiple of M and k is an integer multiple of V;
[0062] Construct a first matrix of size (m / M*d)*(k / (L / N)), and fill each non-zero element of the current model weight matrix into the first matrix in a row-by-row traversal manner to obtain a compressed data matrix corresponding to the current model weight matrix;
[0063] Construct a second matrix of size (m / M*d)*(k / V), and perform filling processing on the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix;
[0064] Construct a third matrix of size (m / M*d)*(k / (L / N)), and fill the third matrix according to the positions of each structured sparse computing storage unit in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing storage unit, so as to obtain a metadata matrix corresponding to the current model weight matrix.
[0065] For the sake of illustration, Figure 2 shows a schematic comparison structure diagram of a model weight matrix applicable to each embodiment of the present invention and a set of matching compression matrices.
[0066] Specifically, as Figure 2 shown, it is a schematic diagram of generating a set of matching compression matrices for a model weight matrix with m = 4 and k = 16. In this model weight matrix, there are 4 compression units with M = 2 and V = 8. In each compression unit, there is d = 1 valid row, and each valid row contains 2 structured sparse computing storage units. Each structured sparse computing storage unit is in a sparse mode with N:L being 2:4. That is, the A, B,..., L filled in each matrix position of this model weight matrix represent non-zero data, and the blank positions represent data filled with 0 values.
[0067] When constructing the compressed data matrix for this model weight matrix, first construct a first matrix of (m / M*d)*(k / (L / N)) being 2*8. Then, in the way of traversing row by row, fill each non-zero data A, B,..., L in the model weight matrix into this first matrix to obtain a matching compressed data matrix.
[0068] Similarly, when constructing the index matrix for this model weight matrix, first construct a second matrix of (m / M*d)*(k / V) being 2*2. Then, fill the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix.
[0069] Specifically, in an optional implementation manner of this embodiment, filling the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix may specifically include:
[0070] Traverse a current compression unit in the current model weight matrix in sequence, and locate d longitudinal matrix positions in the second matrix that match the current compression unit according to the position of the current compression unit in the current model weight matrix;
[0071] Identify the row positions of each valid row in the current compression unit, and correspondingly fill the identified row positions into d vertical matrix positions;
[0072] Return to execute the operation of sequentially traversing a current compression unit in the current model weight matrix until the processing of all compression units in the current model weight matrix is completed, so as to obtain an index matrix corresponding to the current model weight matrix.
[0073] As Figure 2 shown, in this model weight matrix, first traverse a 2*8 compression unit 1 containing A, B, C, and D. Since d = 1, based on the upper left corner position of this compression unit 1 in the model weight matrix, locate 1 corresponding upper left corner vertical matrix position in the second matrix. Then, identify the row positions of the valid rows in this compression unit 1, that is, the rows where A, B, C, and D are located, that is, row 0. Then, fill 0 correspondingly into the upper left corner matrix position in the second matrix. And so on, until the row positions of each valid row in the four compression units in the model weight matrix are correspondingly filled into the second matrix to obtain a matching index matrix.
[0074] Similarly, when constructing the metadata matrix for this model weight matrix, first construct a third matrix with (m / M*d)*(k / (L / N)) being 2*8. Then, according to the positions of each structured sparse computing storage unit in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing storage unit, perform filling processing on the third matrix to obtain a metadata matrix corresponding to the current model weight matrix.
[0075] Correspondingly, in an optional implementation manner of this embodiment, performing filling processing on the third matrix according to the positions of each structured sparse computing storage unit in the current model weight matrix and the positions of each non-zero data in the corresponding structured sparse computing storage unit to obtain a metadata matrix corresponding to the current model weight matrix may include:
[0076] Sequentially traverse a current structured sparse computing storage unit in the current model weight matrix, and locate N horizontal matrix positions matching the current structured sparse computing storage unit in the third matrix according to the position of the current structured sparse computing storage unit in the current model weight matrix;
[0077] Correspondingly fill the column positions where each non-zero data is located in the current structured sparse computing storage unit into the N horizontal matrix positions;
[0078] Return to perform the operation of sequentially traversing a current structured sparse computing storage unit in the current model weight matrix until all the structured sparse computing storage units in the current model weight matrix are processed, so as to obtain a metadata matrix corresponding to the current model weight matrix.
[0079] Continue as Figure 2 shown, the various structured sparse computing storage units included in the model weight matrix can be sequentially traversed by traversing row by row. For example, first, the structured sparse computing storage unit 1 containing non-zero elements A and B is traversed. Since this structured sparse computing storage unit 1 is located in the upper left corner of the model weight matrix, starting from the upper left corner of the third matrix, 2 horizontal matrix positions with N = 2 can be selected. Then, the column positions where A and B are located in the structured sparse computing storage unit 1, A is in the 0th column and B is in the 2nd column, are correspondingly filled into the first two column positions of the first row of this third matrix. And so on, after the positions of the non-zero data in the 2 * 4 structured sparse computing storage units in the above model weight matrix are all correspondingly filled into this third matrix, a matching metadata matrix is obtained.
[0080] It can be understood that by constructing the above compression data matrix, index matrix, and metadata data, the sparsified model weight matrix with the above specific structure can be uniquely determined. Furthermore, the various sparsified model weight matrices included in the mixture of experts model can be efficiently compressed and stored, thereby effectively reducing the storage overhead of the mixture of experts model on the configured computing chip. In addition, when the mixture of experts model deployed in the computing chip performs calculations based on the above compressed form of the model weight matrix, it can give full play to the hardware computing advantages of each sparse computing unit in the computing chip, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0081] Figure 3 The flowchart of another storage optimization method for the mixture of experts model provided by the embodiment of the present invention is based on the above embodiments for refinement. In this embodiment, specifically: after the operation of storing the various sparsified model weight matrices in the target mixture of experts model as matching compressed matrix sets respectively, the operation of calling each sparse computing unit on the computing chip to perform matching sparse calculations based on the target mixture of experts model with optimized storage is further added. And the rearrangement operation of the compressed matrix set and the associated storage operation of the selection array are synchronously increased.
[0082] Correspondingly, as Figure 3 shown, the method includes:
[0083] S310. Load the target mixture of experts model into the computing chip and obtain the various sparsified model weight matrices in the target mixture of experts model.
[0084] Among them, the sparsified model weight matrix includes multiple compression units of a set size. Each compression unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode.
[0085] S320. Generate a set of compression matrices corresponding to each model weight matrix respectively.
[0086] Among them, the set of compression matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix within their respective compression units, and a metadata matrix for storing the positions of non-zero data within their respective structured sparse computing storage units.
[0087] S330. Store each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compression matrices respectively, so as to optimize the storage of the target mixture-of-experts model.
[0088] S340. Obtain a sparse computing unit requirement file that matches the computing chip.
[0089] Among them, the sparse computing unit requirement file defines the data access positions of each thread executed within the sparse computing unit for each model weight matrix in the target mixture-of-experts model during sparse computing.
[0090] S350. Re-arrange the data layout modes of the compressed data matrix, index matrix, and metadata matrix in each set of compression matrices according to the sparse computing unit requirement file.
[0091] S360. Store each re-arranged set of compression matrices in the set memory of the computing chip, so as to achieve continuous storage of data with continuous access characteristics.
[0092] Generally speaking, when loading data, a computing chip (for example, GPU) adopts the technology of Memory Access Coalescing. Its core idea is to merge the memory accesses of multiple threads into a larger memory access operation, which can reduce access latency and improve bandwidth utilization. At the same time, its data loading also has the following characteristics:
[0093] 1. Aligned access: Memory access in a GPU is typically aligned according to a certain stride, usually 32 bytes, 64 bytes or larger. If multiple threads operate on consecutive memory addresses, these accesses can usually be merged into a single memory access operation, thereby reducing the memory access overhead. 2. Merged memory access: For example, if multiple threads access adjacent addresses in memory, the GPU can merge these accesses into a single memory request, avoiding multiple memory accesses, thereby reducing latency and bandwidth waste.
[0094] Based on the above data loading idea, in each embodiment of the present invention, according to the data reading requirements of the sparse computing unit in the computing chip for implementing calculations, the data arrangement methods of the compressed data matrix, the index matrix, and the metadata matrix in each set of compressed matrices are rearranged to further improve the data reading efficiency during subsequent implementation of calculations.
[0095] Specifically, by parsing the sparse computing unit requirement file provided with the computing chip when leaving the factory, the data reading requirements of the sparse computing unit can be obtained. Specifically, in this sparse computing unit requirement file, the data access positions of each thread in the sparse computing unit for each model weight matrix in the target mixture-of-experts model are recorded.
[0096] Among them, the sparse computing unit can be understood as a special hardware circuit in the computing chip for performing matrix multiplication calculations on double sparse matrices. Among them, a double sparse matrix refers to a left operand matrix and a right operand matrix used for matrix multiplication calculations that are both sparse matrices. Generally, a computing chip contains multiple sparse computing units. Inside each sparse computing unit, one or more threads can be started to perform corresponding computing tasks.
[0097] As an example rather than a limitation, in Figure 4 a comparison schematic diagram of a sparse computing unit requirement and the corresponding data rearrangement result applicable to the embodiments of the present invention is shown.
[0098] As Figure 4 shown, in the table on the left in Figure 4 the access positions of each thread adapted in each sparse computing unit to the matrix elements in the matrix for a specific matrix (for example, the compressed data matrix) are recorded. Among them, T0, T1, …, T 29respectively represent threads, and different threads belong to different sparse computing units. T0{3...0} in the 0th row and 0...3rd columns of the left table represents that the data in the 0th row and 0...3rd columns of the compressed data matrix is loaded into the data position {3...0} in the target vector used by T0 thread for calculation. Furthermore, the T0 thread needs to read the data in the 0th row and 0...3rd columns of the compressed data matrix from the storage space where the compressed data matrix is located and store it at the matching position in the target vector.
[0099] Similarly, T0{19...16} in the 8th row and 0...3rd columns of the left table represents that the data in the 8th row and 0...3rd columns of the compressed data matrix is loaded into the data position {19...16} in the target vector used by T0 thread for calculation. Furthermore, the T0 thread needs to read the data in the 8th row and 0...3rd columns of the compressed data matrix from the storage space where the compressed data matrix is located and store it at the matching position in the target vector.
[0100] It can be understood that considering the calculation characteristics of the actual double sparse matrix, each thread does not read data from consecutive positions in a matrix. This data reading method has low efficiency and poor bandwidth utilization, which will reduce the calculation efficiency of the double sparse matrix to a certain extent. Based on this, in various embodiments of the present invention, according to the data reading characteristics of each thread in the sparse computing unit, the data arrangement methods of the compressed data matrix, index matrix, and metadata matrix in each compressed matrix set are rearranged.
[0101] Specifically, as shown in the left and right tables in Figure 4 all 32 calculation data required by the T0 thread are located in the 0...15th columns of the 0th row and the 0...15th columns of the 8th row of the compressed data matrix respectively. However, according to the conventional data storage method, these 32 data are not consecutive. In other words, a thread needs to perform two memory transfers to obtain the data required by the sparse computing unit, which wastes the memory bandwidth of the computing chip.
[0102] To improve memory utilization, as can be seen from the right table in Figure 4 the above 32 non-consecutively stored data can be continuously arranged in a new matrix. Through such optimization, each thread needs to load 32 consecutive bits of data, and the complete data loading operation can be achieved through one memory transfer. At this time, the T0 thread can perform consecutive data reading on the data that needs to be continuously accessed, greatly reducing the memory access overhead and data reading latency, and effectively improving the bandwidth utilization.
[0103] S370. Identify each hybrid model layer included in the target mixture-of-experts model, and obtain the expert assignment pattern corresponding to each hybrid model layer respectively.
[0104] Among them, the expert assignment pattern is used to indicate the row position where the valid row of the intermediate activation value assigned to its own hybrid model layer is located during the implementation of the calculation.
[0105] As mentioned above, the characteristic of the mixture-of-experts model is that it is necessary to determine which expert should be activated and participate in the calculation of the output when processing a specific input. That is, when implementing the calculation based on the target mixture-of-experts model, for any hybrid model layer, its intermediate activation value is assigned by the intermediate routing layer to different experts for processing.
[0106] For any expert, the intermediate activation values not assigned to it for processing thus become structured sparse values. That is, the input to each hybrid model layer is a sparse matrix of intermediate activation values in the form of a sparse matrix, and one or more matrix rows in this sparse matrix of intermediate activation values are sparse rows (all 0).
[0107] The traditional processing method is that for each expert in each hybrid model layer, the data (data in non-sparse rows) of the original sparse matrix of intermediate activation values that is assigned is copied into new data, which will bring additional computational, memory, and bandwidth requirements. Based on this, the embodiments of the present invention further propose a structured sparse representation method based on column vectors to avoid the additional data copying process, thereby greatly reducing the runtime overhead and improving the model running efficiency.
[0108] That is, directly store the sparse matrix of intermediate activation values with sparse characteristics for each hybrid model layer respectively, and avoid the additional cost brought by copying to obtain new data.
[0109] S380. Generate a selection array corresponding to each hybrid model layer according to the expert assignment pattern.
[0110] In this embodiment, since the sparse matrix of intermediate activation values is directly stored corresponding to each hybrid model layer respectively, therefore, in order to meet the computational requirements of the sparse computing unit for the sparse matrix of intermediate activation values, it is necessary to synchronously store the selection array for the sparse matrix of intermediate activation values.
[0111] Among them, the selection array is used to describe which matrix rows in the sparse matrix of intermediate activation values corresponding to each hybrid model layer are non-sparse rows containing valid inputs.
[0112] By way of example and not limitation, in Figure 5 shows a schematic diagram of the control structure of a sparse matrix of intermediate activation values and a matching selection array applicable to the embodiments of the present invention. In Figure 5The specific form of an intermediate activation value sparse matrix adapted by a hybrid model layer is shown in the left matrix. In this intermediate activation value sparse matrix, the 0th row, the 3rd row, and the 5th row are set as valid rows, and the remaining rows are sparse rows of all 0s. At this time, for the expert allocation mode of this hybrid model layer, the intermediate activation values output by the 0th, 3rd, and 5th experts. Furthermore, a selection array in the form of {0, 3, 5} can be constructed to describe the positions of the valid rows in this intermediate activation value sparse matrix.
[0113] S390. Associatively store the selection arrays of each hybrid model layer with the set of compressed matrices of each model weight matrix, so that when performing calculations, the selection array can be combined with the intermediate activation value sparse matrix to be processed and sparse calculations can be performed with the matching set of compressed matrices.
[0114] S3100. Invoke each sparse calculation unit on the computing chip and perform matching sparse calculations based on the target mixture-of-experts model optimized for storage.
[0115] In this embodiment, since the target mixture-of-experts model needs to perform matrix multiplication on the model weight matrix of each model layer and the input activation value matrix input by this model layer during model inference, double sparse matrix multiplication is involved in this calculation process. In each embodiment of the present invention, by compressing and storing the model weight matrix as a set of compressed matrices and synchronously storing the selection array of each hybrid model layer, the sparse calculation units in the computing chip can be fully utilized to improve the calculation efficiency.
[0116] Figure 6 It is a schematic structural diagram of a storage optimization device for a mixture-of-experts model provided by an embodiment of the present invention. As Figure 6 shown, the device includes a model weight matrix acquisition module 610, a compressed matrix set generation module 620, and a storage optimization module 630, where:
[0117] The model weight matrix acquisition module 610 is configured to load the target mixture-of-experts model into the computing chip and acquire various sparsified model weight matrices in the target mixture-of-experts model;
[0118] Among them, the sparsified model weight matrix includes multiple compressed units of set sizes, each compressed unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse calculation storage unit; all the structured sparse calculation storage units are in the same sparse mode;
[0119] The compressed matrix set generation module 620 is configured to generate a set of compressed matrices corresponding to each model weight matrix respectively;
[0120] Among them, the compression matrix set includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in their respective compression units, and a metadata matrix for storing the positions of non-zero data in their respective structured sparse computing storage units;
[0121] A storage optimization module 630 is configured to store each sparsified model weight matrix in the target mixture-of-experts model as a corresponding compression matrix set respectively, so as to optimize the storage of the target mixture-of-experts model.
[0122] The technical solution of the embodiment of the present invention, by converting each sparsified model weight matrix in a specific format in the target mixture-of-experts model into a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in their respective compression units, and a metadata matrix for storing the positions of non-zero data in their respective structured sparse computing storage units, and then performing optimized storage in a computing chip for implementing the calculation of the mixture-of-experts model, can effectively reduce the storage resource overhead of the computing chip. In addition, when using the model weight matrix in the above data compression form for model calculation, the sparse computing hardware in the computing chip can be fully utilized, effectively improving the computing efficiency while reducing the bandwidth overhead.
[0123] Based on the above embodiments, the sparsified model weight matrix may specifically include multiple compression units of size M*V;
[0124] Each compression unit specifically includes d valid rows, where 1≤d<M; and
[0125] Each structured sparse computing storage unit is in an N:L sparse mode, where L is the total number of data included in the structured sparse computing storage unit, and N is the number of non-zero data included in the structured sparse computing storage unit.
[0126] Based on the above embodiments, the compression matrix set generation module 620 may specifically include:
[0127] A current model weight matrix acquisition unit for acquiring a current model weight matrix of size m*k being processed, where m is an integer multiple of M and k is an integer multiple of V;
[0128] A compressed data matrix construction subunit for constructing a first matrix of size (m / M*d)*(k / (L / N)), and filling each non-zero element of the current model weight matrix into the first matrix in a row-by-row traversal manner to obtain a compressed data matrix corresponding to the current model weight matrix;
[0129] An index matrix construction subunit for constructing a second matrix of size (m / M*d) * (k / V), and filling the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in its affiliated compression unit to obtain an index matrix corresponding to the current model weight matrix;
[0130] A metadata matrix construction subunit for constructing a third matrix of size (m / M*d) * (k / (L / N)), and filling the third matrix according to the positions of each structured sparse computing and storage unit in the current model weight matrix and the positions of each non-zero data in its affiliated structured sparse computing and storage unit to obtain a metadata matrix corresponding to the current model weight matrix.
[0131] Based on the above embodiments, the index matrix construction subunit is specifically configured to:
[0132] Traverse a current compression unit in the current model weight matrix in sequence, and locate d longitudinal matrix positions in the second matrix that match the current compression unit according to the position of the current compression unit in the current model weight matrix;
[0133] Identify the row positions of each valid row in the current compression unit, and fill the identified row positions into the d longitudinal matrix positions correspondingly;
[0134] Return and perform the operation of traversing a current compression unit in the current model weight matrix in sequence until the processing of all compression units in the current model weight matrix is completed to obtain an index matrix corresponding to the current model weight matrix.
[0135] Based on the above embodiments, the metadata matrix construction subunit is specifically configured to:
[0136] Traverse a current structured sparse computing and storage unit in the current model weight matrix in sequence, and locate N transverse matrix positions in the third matrix that match the current structured sparse computing and storage unit according to the position of the current structured sparse computing and storage unit in the current model weight matrix;
[0137] Fill the column positions where each non-zero data is located in the current structured sparse computing and storage unit into the N transverse matrix positions correspondingly;
[0138] Return and perform the operation of traversing a current structured sparse computing and storage unit in the current model weight matrix in sequence until the processing of all structured sparse computing and storage units in the current model weight matrix is completed to obtain a metadata matrix corresponding to the current model weight matrix.
[0139] Based on the above embodiments, it may further include a sparse computing module for:
[0140] After storing the sparse model weight matrices in the target mixture-of-experts model as corresponding sets of compressed matrices respectively, call each sparse computing unit on the computing chip to perform corresponding sparse computations based on the target mixture-of-experts model with storage optimization.
[0141] Based on the above embodiments, it may further include a rearrangement module for:
[0142] After storing the sparse model weight matrices in the target mixture-of-experts model as corresponding sets of compressed matrices respectively, obtain a sparse computing unit requirement file matching the computing chip;
[0143] Among them, the sparse computing unit requirement file defines the data access positions of each thread executed in the sparse computing unit to each model weight matrix in the target mixture-of-experts model during sparse computation;
[0144] According to the sparse computing unit requirement file, rearrange the data layout modes of the compressed data matrices, index matrices, and metadata matrices in each set of compressed matrices;
[0145] Store each set of rearranged compressed matrices in the set memory of the computing chip to achieve continuous storage of data with continuous access characteristics.
[0146] Based on the above embodiments, it may further include a selection array construction module for:
[0147] Before calling each sparse computing unit on the computing chip to perform corresponding sparse computations based on the target mixture-of-experts model with storage optimization, identify each mixture model layer included in the target mixture-of-experts model and obtain the expert allocation mode corresponding to each mixture model layer respectively;
[0148] Among them, the expert allocation mode is used to indicate the row position where the valid rows of the intermediate activation values assigned to its own mixture model layer are located during computation;
[0149] Generate a selection array corresponding to each mixture model layer according to the expert allocation mode;
[0150] Associate and store the selection arrays of each mixture model layer with the sets of compressed matrices of each model weight matrix, so that during computation, the selection arrays can be combined with the sparse matrix of intermediate activation values to be processed and perform sparse computations with the corresponding sets of compressed matrices.
[0151] The storage optimization device of the mixture of experts model provided by the embodiments of the present invention can execute the storage optimization method of the mixture of experts model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0152] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0153] Figure 7 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0154] As Figure 7 shown, the electronic device 10 includes at least one computing chip 11, and a memory communicatively connected to the at least one computing chip 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one computing chip. The computing chip 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The computing chip 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0155] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0156] The computing chip 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing chip 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various computing chips running machine learning model algorithms, digital signal computing chips (DSPs), and any suitable computing chips, controllers, microcontrollers, etc. The computing chip 11 executes the various methods and processes described above. For example, it executes the storage optimization method of the mixture-of-experts model as described in any embodiment of the present invention.
[0157] That is, load the target mixture-of-experts model into the computing chip, and obtain each sparsified model weight matrix in the target mixture-of-experts model;
[0158] Among them, the sparsified model weight matrix includes multiple compression units of a set size. Each compression unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing and storage unit; all the structured sparse computing and storage units are in the same sparse mode;
[0159] Generate a set of compression matrices corresponding to each model weight matrix respectively;
[0160] Among them, the set of compression matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of the valid rows in the model weight matrix in the respective compression units, and a metadata matrix for storing the positions of the non-zero data in the respective structured sparse computing and storage units;
[0161] Store each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compression matrices respectively, so as to achieve storage optimization of the target mixture-of-experts model.
[0162] Further, Figure 8 shows a structural diagram of a computing chip applicable to the embodiment of the present invention. As Figure 8 shown, the computing chip further includes: a plurality of sparse computing units 810; the sparse computing units 810 are used to implement sparse computing.
[0163] In some embodiments, the storage optimization method of the mixture-of-experts model as described in any embodiment of the present invention can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the computing chip 11, one or more steps of the storage optimization method of the mixture-of-experts model as described in any embodiment of the present invention above can be executed. Alternatively, in other embodiments, the computing chip 11 can be configured to execute the storage optimization method of the mixture-of-experts model as described in any embodiment of the present invention by any other suitable means (e.g., by means of firmware).
[0164] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable computing chip, which can be a dedicated or general-purpose programmable computing chip, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0165] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the computing chip of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the computer programs are executed by the computing chip, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0166] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0167] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0168] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0169] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0170] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0171] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A storage optimization method for a mixture of experts model, characterized in that Including: Loading the target mixture-of-experts model into a computing chip and obtaining various sparsified model weight matrices in the target mixture-of-experts model; Among them, the sparsified model weight matrix includes multiple compression units of a set size. Each compression unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing storage unit; all structured sparse computing storage units are in the same sparse pattern; a valid row is a matrix row with at least one non-zero data, a sparse row is a matrix row with all data being 0, and the sparse pattern is the proportion of non-zero data in all data in the structured sparse computing storage unit; Generating a set of compression matrices corresponding to each model weight matrix respectively; Among them, the set of compression matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the current compression unit in the model weight matrix, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing storage unit; Storing each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compression matrices respectively to achieve storage optimization of the target mixture-of-experts model; Obtaining a sparse computing unit requirement file matching the computing chip; Among them, the sparse computing unit requirement file defines the data access positions of each thread executed in each sparse computing unit on the computing chip for each model weight matrix in the target mixture-of-experts model; Rearranging the data layout of the compressed data matrix, index matrix, and metadata matrix in each set of compression matrices according to the sparse computing unit requirement file; Storing the rearranged sets of compression matrices in a set memory of the computing chip to achieve continuous storage of data with continuous access characteristics.
2. The method according to claim 1, wherein The sparsified model weight matrix specifically includes multiple compression units of size M*V; Specifically, each compression unit includes d valid rows, where 1≤d<M; and All structured sparse computing storage units are in the sparse pattern of N:L, where L is the total number of data included in the structured sparse computing storage unit, and N is the number of non-zero data included in the structured sparse computing storage unit.
3. The method according to claim 2, wherein Generating a set of compression matrices corresponding to each model weight matrix respectively, including: Obtaining the current model weight matrix of size m*k being processed, where m is an integer multiple of M and k is an integer multiple of V; Constructing a first matrix of size (m / M*d)*(k / (L / N)), and filling each non-zero element of the current model weight matrix into the first matrix row by row to obtain a compressed data matrix corresponding to the current model weight matrix; Constructing a second matrix of size (m / M*d)*(k / V), and filling the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in the corresponding compression unit to obtain an index matrix corresponding to the current model weight matrix; Construct a third matrix of size (m / M*d)*(k / (L / N)), and fill the third matrix according to the positions of each structured sparse computing storage unit in the current model weight matrix and the positions of each non-zero data in its corresponding structured sparse computing storage unit, to obtain a metadata matrix corresponding to the current model weight matrix.
4. The method according to claim 3, wherein Fill the second matrix according to the positions of each compression unit in the current model weight matrix and the positions of each valid row in its corresponding compression unit, to obtain an index matrix corresponding to the current model weight matrix, including: Traverse a current compression unit in the current model weight matrix in sequence, and locate d vertical matrix positions in the second matrix that match the current compression unit according to the position of the current compression unit in the current model weight matrix; Identify the row positions where each valid row is located in the current compression unit, and fill the identified row positions into the d vertical matrix positions correspondingly; Return to execute the operation of traversing a current compression unit in the current model weight matrix in sequence until the processing of all compression units in the current model weight matrix is completed, to obtain an index matrix corresponding to the current model weight matrix.
5. The method according to claim 3, wherein Fill the third matrix according to the positions of each structured sparse computing storage unit in the current model weight matrix and the positions of each non-zero data in its corresponding structured sparse computing storage unit, to obtain a metadata matrix corresponding to the current model weight matrix, including: Traverse a current structured sparse computing storage unit in the current model weight matrix in sequence, and locate N horizontal matrix positions in the third matrix that match the current structured sparse computing storage unit according to the position of the current structured sparse computing storage unit in the current model weight matrix; Fill the column positions where each non-zero data is located in the current structured sparse computing storage unit into the N horizontal matrix positions correspondingly; Return to execute the operation of traversing a current structured sparse computing storage unit in the current model weight matrix in sequence until the processing of all structured sparse computing storage units in the current model weight matrix is completed, to obtain a metadata matrix corresponding to the current model weight matrix.
6. The method according to any one of claims 1-5, characterized in that, After storing each sparsified model weight matrix in the target mixture-of-experts model as a corresponding set of compression matrices respectively, it further includes: Call each sparse computing unit on the computing chip to perform corresponding sparse computing based on the target mixture-of-experts model with storage optimization.
7. The method according to claim 6, wherein Before calling each sparse computing unit on the computing chip to perform corresponding sparse computing based on the target mixture-of-experts model with storage optimization, it further includes: Identify each mixture model layer included in the target mixture-of-experts model, and obtain the expert assignment pattern corresponding to each mixture model layer respectively; Among them, the expert assignment pattern is used to indicate the row position where the valid rows of the intermediate activation values assigned to its own mixture model layer are located during the implementation of the calculation; Generate a selection array corresponding to each mixture model layer respectively according to the expert assignment pattern; Associate and store the selection arrays of each hybrid model layer with the set of compressed matrices of each model weight matrix, so that when performing calculations, the selection arrays can be combined with the sparse matrix of intermediate activation values to be processed and sparse calculations can be performed with the matching set of compressed matrices.
8. A storage optimization device for a mixture of experts model, characterized in that, It includes: A model weight matrix acquisition module, configured to load the target mixture-of-experts model into the computing chip and acquire various sparsified model weight matrices in the target mixture-of-experts model; Among them, the sparsified model weight matrix includes multiple compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse calculation storage unit; all the structured sparse calculation storage units are in the same sparse pattern; a valid row is a matrix row with at least one non-zero data, a sparse row is a matrix row with all data being zero, and the sparse pattern is the proportion of non-zero data in all data in the structured sparse calculation storage unit; A compressed matrix set generation module, configured to generate a set of compressed matrices corresponding to each model weight matrix respectively; Among them, the set of compressed matrices includes a compressed data matrix for storing non-zero data in the model weight matrix, an index matrix for storing the positions of valid rows in the model weight matrix in the corresponding compression unit, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse calculation storage unit; A storage optimization module, configured to store each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compressed matrices respectively to achieve storage optimization of the target mixture-of-experts model; The device further includes: a rearrangement module, configured to: After storing each sparsified model weight matrix in the target mixture-of-experts model as a matching set of compressed matrices respectively, obtain a sparse calculation unit requirement file matching the computing chip; Among them, the sparse calculation unit requirement file defines the data access positions of each thread in each sparse calculation unit on the computing chip to the model weight matrices in the target mixture-of-experts model when performing sparse calculations; According to the sparse calculation unit requirement file, rearrange the data arrangement modes of the compressed data matrix, the index matrix and the metadata matrix in each set of compressed matrices; Store the rearranged sets of compressed matrices in the set memory of the computing chip to achieve continuous storage of data with continuous access characteristics.
9. An electronic device, characterized in that, The electronic device includes: At least one computing chip; and A memory communicatively connected to the at least one computing chip; wherein, The memory stores a computer program executed by the at least one computing chip, and the computer program is executed by the at least one computing chip so that the at least one computing chip can execute the storage optimization method of the mixture-of-experts model according to any one of claims 1-7.
10. The electronic device according to claim 9, wherein The computing chip further includes: a plurality of sparse calculation units; The sparse calculation unit is configured to perform sparse calculations.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computing chip, implement the storage optimization method of the mixture-of-experts model according to any one of claims 1-7.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a computing chip, implements the storage optimization method of the mixture-of-experts model according to any one of claims 1-7.