Method, apparatus, device, medium and product for accelerating multiplication of double sparse matrices

By storing the double sparse matrix in a specific compression format and carrying data in blocks, and using the sparse calculation unit of the computing chip to perform matrix multiplication calculation, the problem of low efficiency of double sparse mode data matrix multiplication calculation in the prior art is solved, and efficient resource utilization and acceleration performance are achieved.

CN119806639BActive Publication Date: 2025-06-20SHANGHAI SUIYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510308459.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-20
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

Existing data flow optimization, calculation and orchestration methods cannot efficiently perform matrix multiplication calculation of double sparse mode data.

Method used

By compressing and storing the sparse matrix in a specific compression format, and transferring it from global memory step by step from global memory in the form of data chunks to the hardware registers of the computing chip, multiplication calculation is performed using the sparse calculation unit of the computing chip.

Benefits of technology

The hardware acceleration performance of sparse computing units in the computing chip is fully utilized, and the overhead of computing, bandwidth and storage resources in double sparse matrix multiplication operations are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806639B_ABST
    Figure CN119806639B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device, medium and product for accelerating the multiplication of double sparse matrices. The method includes: locating a set of compressed matrices matching a first sparse matrix and a second sparse matrix in the global memory of a computing chip according to the sparse matrix multiplication requirement; sequentially transferring the set of compressed matrices and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks; and gradually calculating the multiplication result of the first sparse matrix and the second sparse matrix by a sparse computing unit of the computing chip according to the data loaded in batches in the hardware registers. The technical solution of the embodiments of the present invention can give full play to the hardware acceleration performance of the sparse computing unit in the computing chip, and thus greatly optimize the overhead of computing, bandwidth and storage resources in the process of multiplying double sparse matrices, and is particularly applicable to the model calculation scenario of a mixture-of-experts model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer hardware, and in particular, to a method, device, equipment, medium and product for accelerating the multiplication of double sparse matrices. Background Art

[0002] The operator data flow optimization, calculation and scheduling method is a technical solution that can effectively optimize the utilization rate of computing, bandwidth and storage resources in the hardware computing process.

[0003] The existing data flow optimization, calculation and scheduling methods do not implement the matrix multiplication operator for double sparse mode data (that is, both the left and right operands of the operator are sparse data). Therefore, directly using the existing solutions cannot perform efficient matrix multiplication calculations on double sparse mode data. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device, equipment, medium and product for accelerating the multiplication of double sparse matrices to fully exert the acceleration performance of sparse computing hardware when calculating the multiplication of double sparse matrices.

[0005] According to one aspect of the embodiments of the present invention, there is provided a method for accelerating the multiplication of double sparse matrices, including:

[0006] Locate a set of compressed matrices matching the first sparse matrix and the second sparse matrix in the global memory of the computing chip according to the sparse matrix multiplication requirement;

[0007] Wherein, the first sparse matrix includes multiple compressed units of a set size, each compressed unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in the corresponding compressed units, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse computing storage units;

[0008] Gradually transfer the set of compressed matrices and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks;

[0009] Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second sparse matrix according to the data loaded in batches in the hardware registers.

[0010] According to another aspect of the embodiments of the present invention, there is also provided a device for accelerating the multiplication of double sparse matrices, including:

[0011] A multiplication element positioning module, configured to locate a set of compressed matrices and a second sparse matrix that match a first sparse matrix in the global memory of a computing chip according to the requirements of sparse matrix multiplication;

[0012] Wherein, the first sparse matrix includes multiple compressed units of a set size, each compressed unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in their respective compressed units, and a metadata matrix for storing the positions of non-zero data in their respective structured sparse computing storage units;

[0013] A hierarchical transfer module, configured to hierarchically transfer the set of compressed matrices and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks;

[0014] A multiplication calculation module, configured to gradually calculate the multiplication result of the first sparse matrix and the second sparse matrix through the sparse computing units of the computing chip according to the data loaded in batches in the hardware registers.

[0015] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, which includes:

[0016] At least one computing chip; and

[0017] A memory communicatively connected to the at least one computing chip; wherein,

[0018] The memory stores a computer program executable by the at least one computing chip, and when the computer program is executed by the at least one computing chip, the at least one computing chip is enabled to execute the multiplication acceleration method for double sparse matrices according to any embodiment of the present invention.

[0019] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by a computing chip, the multiplication acceleration method for double sparse matrices according to any embodiment of the present invention is implemented.

[0020] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a computer program, and when the computer program is executed by a computing chip, the steps of the multiplication acceleration method for double sparse matrices according to any embodiment of the present invention are implemented.

[0021] In the technical solution of the embodiment of the present invention, after the double sparse matrix that needs to perform multiplication calculation is compressed and stored in a specific compression format, the double sparse matrix in the specific compression format is gradually transferred from the global memory to the hardware register of the computing chip in the form of data blocks, and through the sparse computing unit of the computing chip, according to the data loaded in batches in the hardware register, the multiplication result of the double sparse matrix is gradually calculated. This implementation method fully considers the sparse characteristics when performing matrix multiplication based on the double sparse matrix. By adopting data block and data transfer technologies adapted to this sparse characteristic, the hardware acceleration performance of the sparse computing unit in the computing chip can be fully utilized, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the double sparse matrix multiplication operation process, and is particularly suitable for the model calculation scenario of the mixture of experts model.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 is a flowchart of a method for accelerating the multiplication of a double sparse matrix provided according to an embodiment of the present invention;

[0025] Figure 2 is a flowchart of another method for accelerating the multiplication of a double sparse matrix provided according to an embodiment of the present invention;

[0026] Figure 3 is a schematic diagram of the control structure of a first sparse matrix and a set of matching compressed matrices applicable to the embodiment of the present invention;

[0027] Figure 4 is a block diagram of the implementation of a data block and a multi-level transfer process applicable to the embodiment of the present invention;

[0028] Figure 5 is a schematic diagram of a register mapping relationship applicable to the embodiment of the present invention;

[0029] Figure 6 is a schematic diagram of the structure of a sparse computing unit and an intermediate register cooperating to perform calculations applicable to the embodiment of the present invention;

[0030] Figure 7 It is a schematic structural diagram of a multiplication acceleration device for a double sparse matrix provided according to an embodiment of the present invention;

[0031] Figure 8 It is a schematic structural diagram of an electronic device for implementing a multiplication acceleration method for a double sparse matrix according to an embodiment of the present invention. Specific embodiments

[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] Figure 1 It is a flowchart of a multiplication acceleration method for a double sparse matrix provided according to an embodiment of the present invention. This embodiment is applicable to the situation of multiplying and accelerating a double sparse matrix through a specific sparse computing unit in a computing chip. This method can be executed by a multiplication acceleration device for a double sparse matrix. The device can be implemented in the form of hardware and / or software and is generally configurable in an electronic device including the computing chip. As Figure 1 shown, the method includes:

[0035] S110. Locate a set of compressed matrices and a second sparse matrix that match the first sparse matrix in the global memory of the computing chip according to the sparse matrix multiplication requirement.

[0036] Optionally, a computing chip can be understood as an integrated circuit for implementing a set computing task (e.g., Internet of Things control, high-performance computing, or mobile computing, etc.). The computing chip can be a general computing chip, such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated computing chip, such as an NPU (Neural Processing Unit). This embodiment does not limit this.

[0037] In this embodiment, the above-mentioned method for accelerating the multiplication of double sparse matrices can be encapsulated in the form of an operator interface. For example, a sparse multiplication operator specifically constructed for implementing the multiplication of double sparse matrices is built. Furthermore, the multiplication acceleration method can be triggered by calling the operator interface of the sparse multiplication operator.

[0038] Specifically, the multiplication of double sparse matrices can be understood as that both the left operand and the right operand for matrix multiplication are sparse matrices. Correspondingly, when a request for calling the interface of the sparse multiplication operator is detected, it can be determined that a sparse matrix multiplication requirement is detected. Furthermore, the identification information of the left operand (hereinafter referred to as the first sparse matrix) and the right operand (hereinafter referred to as the second sparse matrix) that need to perform the multiplication of double sparse matrices can be obtained from the interface call request. Based on the above identification information, the compressed matrix set matching the first sparse matrix and the second sparse matrix can be located in the global memory of the computing chip.

[0039] It can be understood that the multiplication of double sparse matrices is implemented in the computing chip. Furthermore, the first sparse matrix and the second sparse matrix that need to be calculated need to be pre-loaded into the computing chip in advance. To improve the computing speed of the computing chip, the computing data is initially stored in the global memory of the computing chip, and subsequently, the computing data can be transported in blocks to the hardware registers adapted to the hardware computing unit in a step-by-step manner, and the final computing result can be obtained by the hardware computing unit based on the data blocks stored in the hardware registers.

[0040] In this embodiment, in order to achieve the final multiplication acceleration, the sparse format of the first sparse matrix as the left operand is specially defined. At the same time, the storage method of the first sparse matrix in the global memory of the computing chip is also optimized accordingly.

[0041] Among them, the first sparse matrix includes multiple compression units of a set size. Each compression unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse pattern; the compression matrix set includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix within their respective compression units, and a metadata matrix for storing the positions of non-zero data within their respective structured sparse computing storage units.

[0042] Specifically, the first sparse matrix includes an integer number (one or more) of compression units. That is, the size (number of rows * number of columns) of the first sparse matrix can be divided evenly by the size of the compression unit. A compression unit can be understood as a small matrix of a specific size, and this small matrix has at least two matrix rows. Among the above at least two matrix rows, there is at least one valid row and at least one sparse row. A valid row can be understood as a matrix row that includes at least one valid data (non-zero), and a sparse row refers to a matrix row where all data are 0.

[0043] Furthermore, each valid row in the compression unit includes an integer number (one or more) of structured sparse computing storage units. Among them, all the structured sparse computing storage units included in the first sparse matrix are in the same sparse pattern. Generally speaking, for the convenience of calculation, the size of the structured sparse computing storage unit can be adapted to the computing scale of the hardware computing unit in the computing chip.

[0044] The sparse pattern can be understood as the proportion of valid data (non-zero value data) in all data. For example, the sparse pattern can be 2:4 or 4:8, etc. That is, for a 2:4 sparse pattern, the structured sparse computing storage unit altogether includes 4 data, and 2 of these data are non-zero data. And the arrangement positions of the above two non-zero data in the structured sparse computing storage unit are not restricted.

[0045] It can be understood that in the first sparse matrix, there are a large number of 0 value data. If the first sparse matrix is directly stored in the global memory of the computing chip according to the original size of the first sparse matrix, it will cause a waste of a large number of storage units. In addition, directly implementing the subsequent double sparse matrix multiplication based on the first sparse matrix of the original size has low computing efficiency and cannot efficiently use the hardware computing unit in the computing chip, that is, the sparse computing unit.

[0046] In view of this, in this embodiment, for the left operand (the first sparse matrix) of the double sparse multiplication calculation with the above special structure, a novel and efficient data compression method is creatively proposed to solve the above technical problems.

[0047] In this embodiment, through a specific data compression method, the first sparse matrix can be tightly stored in the form of a set of compressed matrices. Specifically, three compressed matrices corresponding to the first sparse matrix are stored in the set of compressed matrices, namely a compressed data matrix, an index matrix, and a metadata matrix.

[0048] Among them, the compressed data matrix is used to store the compressed data of the non-zero data in the first sparse matrix, the index matrix is used to store the positions of the valid rows in the first sparse matrix in the corresponding compression units, and the metadata matrix is used to store the positions of the non-zero data in the corresponding structured sparse computing storage units.

[0049] Obviously, through the above data compression method, the specific positions of each non-zero data in the original first sparse matrix can be determined by three small matrices. This data compression method can also effectively reduce the consumption of storage resources in the computing chip and facilitate the implementation of calculations based on the first sparse matrix.

[0050] In this embodiment, the second sparse matrix may not be compressed and stored. However, considering the actual application requirements of sparse matrix multiplication, for example, when applied in a mixture of experts model, the second sparse matrix generally also has a specific data structure. Typically, one or more matrix rows in the second sparse matrix are sparse rows (all 0).

[0051] Based on this, in the global memory of the computing chip, in addition to storing the second sparse matrix in its entirety, a selection array matching the second sparse matrix is further stored. The selection array is used to describe the positions of the valid rows in the second sparse matrix. In a specific example, assuming that in the second sparse matrix, except for rows 0, 3, and 5, the rest are sparse rows, a selection array in the form of {0, 3, 5} can be constructed and stored in association with the second sparse matrix.

[0052] Correspondingly, when locating the second sparse matrix in the global memory of the computing chip, the selection array matching the second sparse matrix can be located synchronously.

[0053] S120. Transfer the set of compressed matrices and the second sparse matrix to the hardware registers of the computing chip level by level in the form of data blocks from the global memory.

[0054] In an optional implementation manner of this embodiment, the computing chip includes a three-level storage architecture of global memory -> shared memory -> hardware registers. Among them, the above three-level storage is physically closer to the hardware computing units used for computing in the computing chip, especially the hardware registers are closely arranged with the hardware computing units.

[0055] Generally speaking, the closer the storage device is to the computing unit, the faster its data reading and writing speed, but its storage capacity is also smaller. Therefore, for a matrix multiplication of C(M×N)=A(M×K)×B(K×N), the complete data of A and B cannot be completely stored in the shared memory or hardware registers. Instead, it is often necessary to transfer the data in chunks through multiple levels of transfer to the hardware registers in order to enable the hardware computing units in the computing chip to calculate the final complete multiplication result in multiple steps.

[0056] In an optional implementation manner of this embodiment, the data of each matrix in the compressed matrix set can be processed in chunks in combination with the special data structure of the compressed matrix set. At the same time, the data of the second sparse matrix can be processed in chunks in combination with the selection array matching the second sparse matrix to improve the data transfer and subsequent multiplication calculation efficiency.

[0057] S130. Through the sparse computing unit of the computing chip, according to the data loaded in multiple steps in the hardware register, gradually calculate the multiplication result of the first sparse matrix and the second sparse matrix.

[0058] In this embodiment, the sparse computing unit can be understood as a special hardware circuit in the computing chip for performing matrix multiplication calculation on double sparse matrices. Such special hardware circuits can accelerate the double sparse matrix multiplication. Generally, there are multiple sparse computing units in the computing chip. Inside each sparse computing unit, one or more threads can be started to perform corresponding sparse computing tasks.

[0059] By using the cooperation of each sparse computing unit in the computing chip, partial data in the first sparse matrix and the second sparse matrix can be loaded in multiple steps from each hardware register for local calculation, and the local calculation results can be combined or accumulated to finally obtain the multiplication result of the first sparse matrix and the second sparse matrix.

[0060] The technical solution of the embodiment of the present invention, after compressing and storing the double sparse matrices that need to perform multiplication calculation in a specific compression format, transfers the double sparse matrices in the above specific compression format from the global memory to the hardware registers of the computing chip in the form of data chunks, and through the sparse computing unit of the computing chip, according to the data loaded in multiple steps in the hardware register, gradually calculates the multiplication result of the double sparse matrices. This implementation method fully considers the sparse characteristics when performing matrix multiplication calculation based on double sparse matrices. By adopting data chunking and data transfer technologies adapted to this sparse characteristic, the hardware acceleration performance of the sparse computing unit in the computing chip can be fully utilized, thereby greatly optimizing the consumption of computing, bandwidth, and storage resources in the double sparse matrix multiplication operation process, and is particularly suitable for the model calculation scenario of the mixture of experts model.

[0061] Figure 2 FIG. 0 is a flowchart of another method for accelerating the multiplication of double sparse matrices provided by an embodiment of the present invention. This embodiment is optimized based on the above embodiments. In this embodiment, the operation of "transferring the compressed matrix set and the second sparse matrix from the global memory to the hardware register of the computing chip in the form of data blocks" is specifically implemented.

[0062] Correspondingly, as Figure 2 shown, the method may include:

[0063] S210. Locate the compressed matrix set and the second sparse matrix that match the first sparse matrix in the global memory of the computing chip according to the sparse matrix multiplication requirement.

[0064] The inventor found through research that when using a mixture of experts model to calculate tasks (typically, model inference tasks), the matrix multiplication calculation between the model weight matrix of each model layer and the intermediate activation value sparse matrix input to the mixture model layer in the mixture of experts model belongs to double sparse matrix multiplication. Furthermore, the methods of the embodiments of the present invention can be applied to the model calculation scenario of the mixture of experts model.

[0065] In this model calculation scenario, the double sparse matrix multiplication described in the embodiments of the present invention can be implemented by calling a pre-encapsulated sparse multiplication operator based on a specific model weight matrix and a matching intermediate activation value sparse matrix.

[0066] Correspondingly, in an optional implementation manner of the embodiment of the present invention, the first sparse matrix is the model weight matrix in the mixture of experts model, and the second sparse matrix is the intermediate activation value sparse matrix input to the mixture model layer in the mixture of experts model.

[0067] Furthermore, the first sparse matrix may specifically include multiple compression units of M*V size;

[0068] Each compression unit specifically includes d valid rows, where 1≤d<M; and

[0069] Each structured sparse calculation and storage unit is in an N:L sparse mode, where L is the total number of data included in the structured sparse calculation and storage unit, and N is the number of non-zero data included in the structured sparse calculation and storage unit.

[0070] Correspondingly, for the first sparse matrix of m*k size, a compressed data matrix of (m / M*d)*(k / (L / N)) size, an index matrix of (m / M*d)*(k / V) size, and a metadata matrix of (m / M*d)*(k / (L / N)) size can be obtained.

[0071] For the sake of convenience of description, in Figure 3The figure shows a comparison structure diagram of a first sparse matrix and a set of matching compression matrices applicable to various embodiments of the present invention.

[0072] Specifically, as Figure 3 shown is a schematic diagram of a set of compression matrices matching a first sparse matrix with m = 4 and k = 16. In this first sparse matrix, there are 4 compression units with M = 2 and V = 8. In each compression unit, there is d = 1 valid row, and each valid row contains 2 structured sparse calculation and storage units, and each structured sparse calculation and storage unit is in a sparse mode with N:L = 2:4. That is, at each matrix position in this first sparse matrix, the filled A, B, …, L represent non-zero data, and the blank positions represent data filled with 0 values.

[0073] Correspondingly, the size of the compressed data matrix corresponding to the first sparse matrix is (m / M*d) * (k / (L / N)) = 2 * 8. This compressed data matrix sequentially stores each non-zero data A, B, …, L in the first sparse matrix.

[0074] Further, the size of the index matrix corresponding to the first sparse matrix is (m / M*d) * (k / V) = 2 * 2. At different positions in this index matrix, the positions of each valid row in the first sparse matrix in the corresponding compression unit are respectively stored. As Figure 3 shown, in this first sparse matrix, there is a 2 * 8 compression unit 1 containing A, B, C, and D. Since d = 1, and the row position where the valid row in this compression unit 1 is located, that is, the row where A, B, C, and D are located, is the 0th row. Furthermore, the 0 stored at the top-left matrix position in the index matrix can be used to identify the above position relationship.

[0075] Further, the size of the metadata matrix corresponding to the first sparse matrix is (m / M*d) * (k / (L / N)) = 2 * 8. At different positions in this metadata matrix, the positions of each non-zero data in the corresponding structured sparse calculation and storage unit are respectively stored. As Figure 3 shown, the structured sparse calculation and storage unit 1 containing non-zero elements A and B is located in the upper left corner of the first sparse matrix. Furthermore, through the 0 and 2 at the first two column positions of the first row of this metadata matrix, they respectively represent the column positions where A and B are located in the structured sparse calculation and storage unit 1. A is located in the 0th column, and B is located in the 2nd column.

[0076] It can be understood that by constructing the above compression data matrix, index matrix, and metadata, the first sparse matrix with the above specific structure can be uniquely determined. Furthermore, the sparsified model weight matrices included in the mixture of experts model can be efficiently compressed and stored, thereby effectively reducing the storage overhead of the mixture of experts model on the configured computing chip.

[0077] S220. Divide the compression data matrix in the global memory into data blocks, and match the index matrix and the metadata matrix for data block division according to the data correspondence relationship between the compression data matrix, the index matrix, and the metadata matrix.

[0078] In this embodiment, since the compression data matrix stores all non-zero data in the first sparse matrix, the data can be divided into blocks according to the data scale of the compression data matrix. During the process of dividing the compression data matrix into data blocks, in order to further consider the efficiency of subsequent data transfer and multiplication calculation, it can be set that during the process of dividing the compression data matrix into data blocks in the global memory, V (the number of columns included in each compression unit) is an integer multiple of the segmentation size in the column segmentation direction.

[0079] For example, assume that in the global memory, when dividing the compression data matrix into multiple m b *k b sized data blocks, it is required that V is an integer multiple of k b , for example, k b is V / 2 or V / 4, etc.

[0080] As shown above, since there is a data or position correspondence relationship among the compression data matrix, the index matrix, and the metadata matrix. After determining the compression data matrix blocks segmented from the compression data matrix, the corresponding index matrix blocks and metadata matrix blocks can be uniquely determined. That is, the index matrix and the metadata matrix are matched for data block division according to the data correspondence relationship between the compression data matrix, the index matrix, and the metadata matrix.

[0081] S230. Divide the second sparse matrix in the global memory according to the selection array.

[0082] As shown above, the second sparse matrix may contain a large number of sparse rows. If these sparse rows are also divided into data blocks and sent to the hardware computing unit for matrix multiplication calculation, it will bring redundant data loading and calculation, reducing the calculation efficiency.

[0083] Based on this, in this embodiment, in the global memory, data chunking can be performed only on the valid data rows in the second sparse matrix based on the valid row positions defined by the selection array. Optionally, one or more valid data rows can be taken as a data chunk each time, or a set number of data in one valid data row can be taken as a data chunk each time, etc.

[0084] S240. Load each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk from the global memory into the shared memory sequentially.

[0085] After data chunking for the set of compressed matrices is completed in the global memory, each of the above-mentioned compressed data matrix chunks, each index matrix chunk, and each metadata matrix chunk can be loaded from the global memory into the shared memory sequentially.

[0086] Based on the above embodiments, in order to prevent data in the shared memory from being restricted by bank conflicts where the same memory bank in the computing chip cannot be accessed by multiple threads simultaneously during access (read / write), in the embodiments of the present invention, data rearrangement in the shared memory is further considered using a specific data rearrangement function based on data offsets.

[0087] Correspondingly, in an optional implementation manner of this embodiment, after each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk are sequentially transferred from the global memory to the shared memory, it may further include:

[0088] Use a preset data rearrangement function to rearrange each compressed data matrix chunk, each index matrix chunk, and each metadata matrix chunk in the shared memory to avoid bank conflicts.

[0089] Specifically, according to the specific type and model of the computing chip, a matching data rearrangement function can be selected from the adapted function library to rearrange the data in the shared memory, and this embodiment does not limit this.

[0090] S250. Transfer each second sparse matrix chunk from the global memory to the shared memory sequentially.

[0091] Similarly, after data chunking of the second sparse matrix is completed in the global memory, each second sparse matrix chunk can be correspondingly transferred from the global memory to the shared memory sequentially.

[0092] Meanwhile, in order to avoid the problem of bank conflicts, in an optional implementation manner of this embodiment, after each second sparse matrix chunk is sequentially transferred from the global memory to the shared memory, it may further include:

[0093] Use a preset data rearrangement function to rearrange each second sparse matrix block in shared memory to avoid bank conflicts.

[0094] Of course, it can be understood that in addition to using the data rearrangement function to rearrange each compressed data matrix block, each index matrix block, each metadata matrix block, and each second sparse matrix block in shared memory, it is also possible to add data padding to each compressed data matrix block, each index matrix block, each metadata matrix block, and each second sparse matrix block in shared memory to avoid bank conflicts.

[0095] S260. Perform secondary partitioning on each compressed data matrix block, each index matrix block, and each metadata matrix block in shared memory.

[0096] In this embodiment, in order to adapt to the storage limitations of hardware registers, secondary partitioning can be performed on each compressed data matrix block, each index matrix block, and each metadata matrix block in shared memory.

[0097] Furthermore, when performing secondary partitioning on each compressed data matrix block, each index matrix block, and each metadata matrix block in shared memory, the secondary partitioning of the compressed data matrix block can also be performed first. During the secondary partitioning of the compressed data matrix block, in order to further consider the efficiency of subsequent data transfer and multiplication calculations, it can be set that during the secondary partitioning of the compressed data matrix block in shared memory, V is also an integer multiple of the segmentation size in the column segmentation direction. For example, secondary segmentation can be selected in the column segmentation direction or no secondary segmentation can be selected (only segmentation in the row direction).

[0098] For example, assume that in global memory, when the compressed data matrix block is secondarily segmented into multiple new data blocks of size m b1 *k b1 , it is required that V is an integer multiple of k b1 , for example, k b1 is V / 4 or V / 8, etc.

[0099] Similarly, since there is a data or position correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix, after the secondary partitioning of the compressed data matrix block is completed, the secondary partitioning of the index matrix block and the metadata matrix block can be performed accordingly.

[0100] S270. Perform secondary partitioning on each second sparse matrix block in shared memory.

[0101] S280. According to the register mapping relationship matching the sparse computing unit, at least one of the secondary blockings of each compressed data matrix, the secondary blockings of each index matrix, and the secondary blockings of each metadata matrix is successively transferred from the shared memory to the hardware register.

[0102] For a more intuitive understanding, in Figure 4 a block diagram showing an implementation of a data blocking and a multi-level transfer process applicable to an embodiment of the present invention is shown. As Figure 4 shown, it describes that after the compressed data matrix A in the first sparse matrix and the second sparse matrix B are respectively blocked in the global memory and the shared memory, and are successively transferred in the order of global memory -> shared memory -> register, after the calculation is completed by the sparse computing unit and gradually written back, the corresponding multiplication result matrix C is obtained in the global memory.

[0103] In this example, an example of data blocking and data transfer of a matrix multiplication of any size C(M×N)=A(M×K)×B(K×N) is given. In this embodiment, in order to implement matrix multiplication calculation, the second sparse matrix B is row-column interchanged, that is, an array is selected to describe the positions of the valid data columns in the second sparse matrix B.

[0104] Specifically, in the global memory, since only the positions of the valid data columns (rows) are included in the selected array, the total amount of data in the selected array is len d pieces, and then the number of valid values in the n dimension of the second sparse matrix B is len d pieces. The compressed data matrix A is evenly divided into several pieces in the m dimension with m b as the size. The second sparse matrix B is divided in the n dimension, and the number of valid column vectors in each block data is the same as the data block divided by the size of n b in the selected array. Considering the storage size limitation of the computing device, the compressed data matrix A and the second sparse matrix B will be further evenly divided in the k direction with k b as the size. The multiplication results of multiple sub-blocks in the k direction are accumulated to obtain the final multiplication result matrix C block result. Based on this data partitioning strategy, the size of the multiplication result matrix C calculated for each block data is m b ×n b , and the result calculated by the block calculation is a part of the final result of the multiplication result matrix C, corresponding to the same row offset of the data block of the compressed data matrix A in the compressed data matrix A and the position where the data block of the second sparse matrix B has the same column offset in the selected array. The parallelism of the double sparse matrix multiplication operator is improved by having each thread group be responsible for one block.

[0105] Similarly, in shared memory, it also involves the process of re - chunking each data chunk transferred from global memory, moving it to registers for storage, and then performing matrix multiplication calculations by an adapted sparse computing unit. This process will not be elaborated here. In this embodiment, the matrix specification of the compressed data matrix A processed by the sparse computing unit once is mi * ki, and the matrix specification of the second sparse matrix B is ki * ni. For more convenient subsequent calculations, the aforementioned V is an integer multiple of ki.

[0106] In an alternative embodiment of this embodiment, the re - chunking of each compressed data matrix, the re - chunking of each index matrix, and the re - chunking of each metadata matrix can all be moved to the register for calculation. Or, considering that register resources are very precious, only the re - chunking of each compressed data matrix and the re - chunking of each metadata matrix can be moved to the register for calculation, while the index matrix is retained in shared memory and only fetched from shared memory for use when needed.

[0107] It should be emphasized that, different from the calculation logic of general hardware computing units, when using a sparse computing unit to accelerate multiplication calculations, the data required for calculation needs to be placed in an adapted instruction register according to the requirements of the sparse instruction set. Therefore, in each embodiment of the present invention, the data in shared memory needs to be moved to a matching hardware register according to the register mapping relationship matching the sparse computing unit.

[0108] Optionally, the register mapping relationship can be read from the specification file configured at the factory of the computing chip. The register hardware file describes the mapping relationship between the calculation data at different positions and the hardware registers.

[0109] Specifically, in Figure 5 shows a schematic diagram of a register mapping relationship applicable to the embodiments of the present invention, and this register mapping relationship is adapted to the re - chunking of the compressed data matrix. As Figure 5 shown, T0{a0, a1} at the upper - left corner position in this register mapping relationship represents: the data in the 0th and 1st columns of the 0th row in the re - chunking of the compressed data matrix are allocated to the register a0 and register a1 matching the thread T0. Similarly, T 0…3 {a4, a5} at the last position in the 0th row represents: the data in the 8th - 15th columns of the 0th row in the re - chunking of the compressed data matrix are respectively allocated to the register a4 and register a5 matching the threads T0, T1, T2, and T3. Among them, different threads T are pre - assigned to different sparse computing units for use.

[0110] Similarly, for the second-level block division of the metadata matrix and the second-level block division of the index matrix, the sparse computing unit also has a corresponding register mapping relationship. Based on the above register mapping relationship, each computing data can be sequentially transferred from the shared memory to the hardware register.

[0111] In an optional implementation manner of this embodiment, according to the register mapping relationship matching the sparse computing unit, a preset hardware instruction can be called to sequentially transfer each compressed data matrix second-level block and each metadata matrix second-level block, or each compressed data matrix second-level block, each index matrix second-level block, and each metadata matrix second-level block, from the shared memory to the hardware register to achieve hardware acceleration of data loading.

[0112] Among them, the hardware instruction is associated with the specifications of the computing chip and can be queried and obtained from the instruction library of the computing chip. For example, it can be a 1dmatrix instruction, etc.

[0113] S290: According to the register mapping relationship matching the sparse computing unit, sequentially load each second sparse matrix second-level block from the shared memory to the hardware register.

[0114] As shown before, by querying and obtaining the register mapping relationship set by the sparse computing unit for the second sparse matrix second-level block, each second sparse matrix second-level block can be sequentially loaded from the shared memory to the hardware register.

[0115] Similarly, in an optional implementation manner of this embodiment, according to the register mapping relationship matching the sparse computing unit, a preset hardware instruction can be called to sequentially load each second sparse matrix second-level block from the shared memory to the hardware register to achieve hardware acceleration of data loading.

[0116] S2100: Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second sparse matrix according to the data loaded in batches in the hardware register.

[0117] In an optional implementation manner of this embodiment, through the sparse computing unit of the computing chip, gradually calculating the multiplication result of the first sparse matrix and the second sparse matrix according to the data loaded in batches in the hardware register may include:

[0118] Through the sparse computing unit, obtain the current compressed data matrix second-level block, the current metadata matrix second-level block, and the current second sparse matrix second-level block currently loaded in the hardware register;

[0119] Through the sparse computing unit, generate a first-level sparse block matrix according to the current metadata matrix second-level block and the current compressed data matrix second-level block;

[0120] Through the sparse computing unit, perform multiplication calculation according to a primary sparse partitioned matrix and the current secondary partitioning of the second sparse matrix.

[0121] As described above, when the first sparse matrix is compressed and stored, the compressed data matrix only stores the non-zero data in each valid row. In the actual first sparse matrix, each valid row contains at least one structured sparse computing storage unit in a set sparse pattern. To ensure the accuracy of the multiplication calculation, it is necessary to first restore the current compressed data matrix secondary partition to the form of the aforementioned structured sparse computing storage unit. Since the metadata matrix stores the positions of non-zero data in the respective structured sparse computing storage units, furthermore, a primary sparse partitioned matrix can be generated according to the current metadata matrix secondary partition and the current compressed data matrix secondary partition.

[0122] That is to say, the primary sparse partitioned matrix can be understood as the matrix after restoring the structured sparse computing storage unit in the current compressed data matrix secondary partition. In a specific example, if the current compressed data matrix secondary partition is {A, B}, and the corresponding current metadata matrix secondary partition is {0, 2}, then a primary sparse partitioned matrix in the form of {A, 0, B, 0} can be restored.

[0123] Among them, through the sparse computing unit, perform multiplication calculation according to the primary sparse partitioned matrix and the current secondary partitioning of the second sparse matrix, which may specifically include:

[0124] Through the sparse computing unit, store the intermediate calculation amount generated during the multiplication calculation into the intermediate register of the computing chip; and through the sparse computing unit, perform the matching multiplication calculation by obtaining the intermediate calculation amount from the intermediate register.

[0125] In the prior art, when using a general computing unit to perform multiplication calculation, the results of matrix multiplication calculation can be directly accumulated, but in essence, it increases the calculation of many redundant data. In contrast, when the sparse computing units of the embodiments of the present invention perform sparse multiplication calculation, when the multiplication calculation iterates along the K direction, since the sparse data comes from different rows, to ensure correctness, when different sub-blocks move along the K direction, the output results need to be mapped to different rows. Traditionally, passing the output results to specific registers according to the index may cause the output C matrix to be reloaded into the local memory, which has a significant impact on the performance of the operator.

[0126] To avoid the above problems, the embodiments of the present invention reduce the frequent memory transfer between the global memory and the register by introducing an additional intermediate register C IR in this way. Among them, Figure 6A schematic diagram of the structure of a sparse computing unit cooperating with an intermediate register to perform calculations applicable to an embodiment of the present invention is shown. As Figure 6 shown, when the sparse computing unit performs sparse matrix multiplication calculations on the current compressed data matrix secondary block ( Figure 6 A in), the current metadata matrix secondary block ( Figure 6 metadata in), and the previous second sparse matrix secondary block ( Figure 6 B in), for multiple result data, during the calculation process, the required intermediate results are loaded into C IR and used for calculation, and after the calculation is completed, the intermediate results are stored back to the corresponding result data position C.

[0127] By adding a new intermediate register to the computing chip, which is a simple hardware improvement, during the process of the sparse computing unit performing double sparse matrix multiplication calculations, the computing performance can be greatly improved and the computing efficiency can be increased.

[0128] Based on the above embodiments, after performing multiplication calculations through the sparse computing unit according to the first sparse block matrix and the current second sparse matrix secondary block, it may further include:

[0129] After determining that the complete multiplication calculation result is stored in the intermediate register, according to the current index matrix secondary block matching the current compressed data matrix secondary block, convert the complete multiplication calculation result into a secondary sparse block matrix; write the secondary sparse block matrix back to the global memory of the computing chip level by level.

[0130] Furthermore, since when obtaining the complete multiplication calculation result, the sparse rows in the first sparse matrix are not considered, therefore, after obtaining the complete multiplication calculation result, in combination with the current index matrix secondary block matching the current compressed data matrix secondary block, add the matching sparse rows to the complete multiplication calculation result to obtain a secondary sparse block matrix as the real block calculation result. As mentioned above, the current index matrix secondary block can also be transferred to the hardware register or only stored in the shared memory, and this embodiment does not limit this.

[0131] The technical solution of the embodiment of the present invention fully considers the sparse characteristics when performing matrix multiplication calculations based on double sparse matrices. By adopting data block and data transfer technologies adapted to this sparse characteristic, the hardware acceleration performance of the sparse computing unit in the computing chip can be fully exerted, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the process of double sparse matrix multiplication operations, and is particularly applicable to the model calculation scenarios of the mixture of experts model.

[0132] Figure 7The figure is a schematic structural diagram of a multiplication acceleration device for a double sparse matrix provided by an embodiment of the present invention. As Figure 7 shown, the device includes: a multiplication element positioning module 710, a step-by-step transfer module 720, and a multiplication calculation module 730, where:

[0133] The multiplication element positioning module 710 is configured to locate a compressed matrix set and a second sparse matrix that match the first sparse matrix in the global memory of the computing chip according to the requirements of sparse matrix multiplication.

[0134] Among them, the first sparse matrix includes multiple compression units of a set size, each compression unit includes at least one valid row and at least one sparse row, and each valid row includes at least one structured sparse calculation storage unit; all the structured sparse calculation storage units are in the same sparse mode; the compressed matrix set includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in the corresponding compression units, and a metadata matrix for storing the positions of non-zero data in the corresponding structured sparse calculation storage units.

[0135] The step-by-step transfer module 720 is configured to step by step transfer the compressed matrix set and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks.

[0136] The multiplication calculation module 730 is configured to gradually calculate the multiplication result of the first sparse matrix and the second sparse matrix through the sparse calculation unit of the computing chip according to the data loaded in batches in the hardware registers.

[0137] The technical solution of the embodiment of the present invention, after compressing and storing the double sparse matrix that needs to perform multiplication calculation in a specific compression format, transfers the double sparse matrix in the above specific compression format from the global memory to the hardware registers of the computing chip in the form of data blocks step by step, and through the sparse calculation unit of the computing chip, gradually calculates the multiplication result of the double sparse matrix according to the data loaded in batches in the hardware registers. This implementation method fully considers the sparse characteristics when performing matrix multiplication calculation based on the double sparse matrix. By adopting data block and data transfer technologies adapted to this sparse characteristic, the hardware acceleration performance of the sparse calculation unit in the computing chip can be fully utilized, thereby greatly optimizing the computing, bandwidth, and storage resource overheads in the double sparse matrix multiplication operation process, and is particularly suitable for the model calculation scenario of the mixture of experts model.

[0138] Based on the above embodiments, the step-by-step transfer module 720 may specifically include:

[0139] The first - type first - level block unit is used to block the compressed data matrix in the global memory and, according to the data correspondence relationship among the compressed data matrix, the index matrix, and the metadata matrix, perform matching data blocks on the index matrix and the metadata matrix;

[0140] The first - type first - level transfer unit is used to successively load each block of the compressed data matrix, each block of the index matrix, and each block of the metadata matrix from the global memory to the shared memory;

[0141] The first - type second - level block unit is used to perform secondary blocking on each block of the compressed data matrix, each block of the index matrix, and each block of the metadata matrix in the shared memory;

[0142] The first - type second - level transfer unit is used to successively transfer at least one of each secondary block of the compressed data matrix, each secondary block of the index matrix, and each secondary block of the metadata matrix from the shared memory to the hardware registers according to the register mapping relationship matching the sparse computing unit.

[0143] Based on the above - mentioned embodiments, it may further include:

[0144] The selection array positioning module is used to locate the compressed matrix set matching the first sparse matrix and the second sparse matrix in the global memory of the computing chip, and at the same time, locate the selection array matching the second sparse matrix in the computing chip, where the selection array is used to describe the valid row positions in the second sparse matrix;

[0145] Correspondingly, the hierarchical transfer module 720 may further include:

[0146] The second - type first - level block unit is used to block the second sparse matrix in the global memory according to the selection array;

[0147] The second - type first - level transfer unit is used to successively transfer each block of the second sparse matrix from the global memory to the shared memory;

[0148] The second - type second - level block unit is used to perform secondary blocking on each block of the second sparse matrix in the shared memory;

[0149] The second - type second - level transfer unit is used to successively load each secondary block of the second sparse matrix from the shared memory to the hardware registers according to the register mapping relationship matching the sparse computing unit.

[0150] Based on the above - mentioned embodiments, it may further include the first - type rearrangement unit, which is used to:

[0151] After each compressed data matrix block, each index matrix block, and each metadata matrix block are successively transferred from global memory to shared memory, a preset data rearrangement function is used to rearrange each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory to avoid bank conflicts; and

[0152] It may further include a second type of rearrangement unit for:

[0153] After each second sparse matrix block is successively transferred from global memory to shared memory, a preset data rearrangement function is used to rearrange each second sparse matrix block in the shared memory to avoid bank conflicts.

[0154] Based on the above embodiments, the first type of secondary transfer unit may specifically be used for:

[0155] According to the register mapping relationship matching the sparse computing unit, a preset hardware instruction is called to successively transfer each second-level block of the compressed data matrix and each second-level block of the metadata matrix, or each second-level block of the compressed data matrix, each second-level block of the index matrix, and each second-level block of the metadata matrix from the shared memory to the hardware registers to achieve hardware acceleration of data loading; and

[0156] The second type of secondary transfer unit may specifically be used for:

[0157] According to the register mapping relationship matching the sparse computing unit, a preset hardware instruction is called to successively load each second-level block of the second sparse matrix from the shared memory to the hardware registers to achieve hardware acceleration of data loading.

[0158] Based on the above embodiments, the multiplication calculation module 730 may specifically include:

[0159] A multiplication calculation element acquisition unit for acquiring the current second-level block of the compressed data matrix, the current second-level block of the metadata matrix, and the current second-level block of the second sparse matrix currently loaded in the hardware registers through the sparse computing unit;

[0160] A first-level sparse block matrix generation unit for generating a first-level sparse block matrix according to the current second-level block of the metadata matrix and the current second-level block of the compressed data matrix through the sparse computing unit;

[0161] A multiplication calculation execution unit for performing multiplication calculation according to the first-level sparse block matrix and the current second-level block of the second sparse matrix through the sparse computing unit.

[0162] Based on the above embodiments, the multiplication calculation execution unit may specifically be used for:

[0163] Through a sparse computing unit, storing the intermediate computation amounts generated during the multiplication computation into an intermediate register of the computing chip; and

[0164] Through the sparse computing unit, by obtaining the intermediate computation amounts from the intermediate register, performing a matching multiplication computation.

[0165] Based on the above embodiments, it may further include a data write-back module for:

[0166] After performing the multiplication computation through the sparse computing unit according to a primary sparse block matrix and a current secondary block of the second sparse matrix, after determining that the complete multiplication computation result is stored in the intermediate register, according to the current secondary block of the index matrix that matches the current compressed data matrix secondary block, converting the complete multiplication computation result into a secondary sparse block matrix;

[0167] Writing back the secondary sparse block matrix to the global memory of the computing chip level by level.

[0168] Based on the above embodiments, the primary sparse matrix may specifically include multiple compression units of size M*V;

[0169] Each compression unit specifically includes d valid rows, where 1≤d<M; and

[0170] Each structured sparse computing storage unit is in a sparse mode of N:L, where L is the total number of data included in the structured sparse computing storage unit, and N is the number of non-zero data included in the structured sparse computing storage unit.

[0171] Based on the above embodiments, during each data block division process, V is an integer multiple of the division size in the column division direction; and

[0172] The matrix specification of the primary sparse matrix processed once in each sparse computing unit is mi*ki, where V is an integer multiple of ki.

[0173] Based on the above embodiments, the primary sparse matrix is the model weight matrix in the mixture-of-experts model, and the second sparse matrix is the intermediate activation value sparse matrix input to the mixture model layer in the mixture-of-experts model.

[0174] The multiplication acceleration device for double sparse matrices provided by the embodiments of the present invention can execute the multiplication acceleration method for double sparse matrices provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0175] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0176] Figure 8 FIG. 1 shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0177] As Figure 8 shown, the electronic device 10 includes at least one computing chip 11, and a memory communicatively connected to the at least one computing chip 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one computing chip. The computing chip 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The computing chip 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0178] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0179] The computing chip 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing chip 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing chips running machine learning model algorithms, digital signal computing chips (DSPs), and any appropriate computing chips, controllers, microcontrollers, etc. The computing chip 11 executes the various methods and processes described above, such as executing the multiplication acceleration method for a double sparse matrix as described in any one of the embodiments of the present invention.

[0180] That is: according to the requirements of sparse matrix multiplication, locate the set of compressed matrices and the second sparse matrix that match the first sparse matrix in the global memory of the computing chip;

[0181] Among them, the first sparse matrix includes multiple compressed units of a set size. Each compressed unit includes at least one valid row and at least one sparse row. Each valid row includes at least one structured sparse computing storage unit; all the structured sparse computing storage units are in the same sparse mode; the set of compressed matrices includes a compressed data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the positions of valid rows in the first sparse matrix in their respective compressed units, and a metadata matrix for storing the positions of non-zero data in their respective structured sparse computing storage units;

[0182] Gradually transfer the set of compressed matrices and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks;

[0183] Through the sparse computing unit of the computing chip, gradually calculate the multiplication result of the first sparse matrix and the second sparse matrix according to the data loaded in batches in the hardware registers.

[0184] In some embodiments, the method for accelerating the multiplication of double sparse matrices as described in any one of the embodiments of the present invention can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the computing chip 11, one or more steps of the method for accelerating the multiplication of double sparse matrices as described in any one of the embodiments of the present invention above can be executed. Alternatively, in other embodiments, the computing chip 11 can be configured to execute the method for accelerating the multiplication of double sparse matrices as described in any one of the embodiments of the present invention by any other suitable means (for example, by means of firmware).

[0185] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable computing chip, which can be a special-purpose or general-purpose programmable computing chip, receiving data and instructions from, and transmitting data and instructions to, a storage system, at least one input device, and at least one output device.

[0186] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the computing chips of a general purpose computer, special purpose computer, or other programmable data processing device, such that the computer programs, when executed by the computing chips, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0187] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0188] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0189] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0190] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0191] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0192] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for accelerating the multiplication of double sparse matrices, characterized in that: include: According to the sparse matrix multiplication requirement, a compressed matrix set and a second sparse matrix matching the first sparse matrix are located in the global memory of the computing chip; Wherein, the first sparse matrix includes a plurality of compression units of set size, each compression unit includes at least one valid row and at least one sparse row, each valid row includes at least one structured sparse computing storage unit; each structured sparse computing storage unit is in the same sparse mode; the compression matrix set includes a compression data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the position of the valid rows in the first sparse matrix in the corresponding compression unit, and a metadata matrix for storing the position of the non-zero data in the corresponding structured sparse computing storage unit; the sparse rows are the matrix rows in which all data are 0, and the sparse mode is the proportion of non-zero data in the structured sparse computing storage unit to all data; The compressed matrix set and the second sparse matrix are moved from the global memory to the hardware registers of the computing chip in the form of data blocks; The multiplication results of the first sparse matrix and the second sparse matrix are gradually calculated by the sparse computing unit of the computing chip according to the data loaded in batches in the hardware register.

2. The method according to claim 1, characterized in that The compressed matrix set is moved from the global memory to the hardware registers of the computing chip in the form of data blocks, including: The compressed data matrix in the global memory is divided into data blocks, and according to the data correspondence between the compressed data matrix, the index matrix and the metadata matrix, the index matrix and the metadata matrix are matched into data blocks; Load each compressed data matrix block, each index matrix block, and each metadata matrix block from the global memory to the shared memory one by one; Secondarily partition each compressed data matrix block, each index matrix block, and each metadata matrix block in the shared memory; According to the register mapping relationship matched with the sparse computing unit, at least one item of each compressed data matrix secondary block, each index matrix secondary block and each metadata matrix secondary block is moved from the shared memory to the hardware register one by one.

3. The method according to claim 2, characterized in that While locating the compressed matrix set and the second sparse matrix matching the first sparse matrix in the global memory of the computing chip, the method further includes: Locating a selection array matching the second sparse matrix in the computing chip, wherein the selection array is used to describe valid row positions in the second sparse matrix; Accordingly, the second sparse matrix is ​​moved from the global memory to the hardware register of the computing chip in the form of data blocks, step by step, including: According to the selection array, the second sparse matrix in the global memory is divided into blocks; Divide each second sparse matrix into blocks and move them from the global memory to the shared memory one by one; Perform secondary block division on each second sparse matrix block in the shared memory; According to the register mapping relationship matched with the sparse computing unit, each second sparse matrix is ​​divided into blocks twice and loaded from the shared memory into the hardware register one by one.

4. The method according to claim 3, characterized in that After each compressed data matrix block, each index matrix block, and each metadata matrix block are successively moved from the global memory to the shared memory, the following further includes: Using a preset data rearrangement function, rearrange each compressed data matrix block, each index matrix block, and each metadata matrix block in a shared memory to avoid storage bank conflicts; and After each second sparse matrix is ​​divided into blocks and moved from the global memory to the shared memory one by one, the method further includes: Using a preset data rearrangement function, each second sparse matrix block is rearranged in a shared memory to avoid storage bank conflicts.

5. The method according to claim 2, characterized in that According to the register mapping relationship matched with the sparse computing unit, at least one item of each compressed data matrix secondary block, each index matrix secondary block and each metadata matrix secondary block is moved from the shared memory to the hardware register one by one, including: According to the register mapping relationship matched with the sparse computing unit, the preset hardware instructions are called to move each compressed data matrix secondary block and each metadata matrix secondary block, or each compressed data matrix secondary block, each index matrix secondary block and each metadata matrix secondary block from the shared memory to the hardware register one by one, so as to realize the hardware acceleration of data loading; and According to the register mapping relationship matching the sparse computing unit, each second sparse matrix is ​​divided into blocks and loaded from the shared memory into the hardware register one by one, including: According to the register mapping relationship matching the sparse computing unit, the preset hardware instructions are called to divide each second sparse matrix into blocks and load them from the shared memory to the hardware registers one by one to achieve hardware acceleration of data loading.

6. The method according to claim 3, characterized in that The multiplication results of the first sparse matrix and the second sparse matrix are calculated step by step according to the data loaded in the hardware register in batches by the sparse computing unit of the computing chip, including: Obtaining, through the sparse computing unit, the current compressed data matrix secondary block, the current metadata matrix secondary block, and the current second sparse matrix secondary block currently loaded in the hardware register; Generate a primary sparse block matrix according to the secondary block of the current metadata matrix and the secondary block of the current compressed data matrix through a sparse computing unit; The multiplication calculation is performed by the sparse computing unit according to the primary sparse block matrix and the secondary block of the current second sparse matrix.

7. The method according to claim 6, characterized in that The sparse computing unit performs a multiplication calculation according to the primary sparse block matrix and the secondary block of the current second sparse matrix, specifically including: The intermediate calculation amount generated in the multiplication calculation process is stored in the intermediate register of the computing chip through the sparse computing unit; and The sparse computing unit performs the matching multiplication calculation by acquiring the intermediate calculation amount from the intermediate register.

8. The method according to claim 7, characterized in that After performing multiplication calculation according to the primary sparse block matrix and the current second sparse matrix secondary block through the sparse computing unit, the method further includes: After determining that the intermediate register stores the complete multiplication calculation result, converting the complete multiplication calculation result into a secondary sparse block matrix according to the current index matrix secondary block matching the current compressed data matrix secondary block; The secondary sparse block matrix is ​​written back to the global memory of the computing chip level by level.

9. The method according to any one of claims 1 to 8, characterized in that The first sparse matrix specifically includes a plurality of compression units of size M*V; Specifically, each compression unit includes d valid rows, where 1≤d<M; and Each structured sparse computing storage unit is in a sparse mode of N:L, where L is the total amount of data contained in the structured sparse computing storage unit, and N is the amount of non-zero data contained in the structured sparse computing storage unit.

10. The method according to claim 9, characterized in that In each data segmentation process, V is an integer multiple of the segmentation size in the column segmentation direction; and The matrix size of the first sparse matrix processed once in each sparse computing unit is mi*ki, where V is an integer multiple of ki.

11. The method according to any one of claims 1 to 8, characterized in that: The first sparse matrix is ​​a model weight matrix in the hybrid expert model, and the second sparse matrix is ​​an intermediate activation value sparse matrix input to the hybrid model layer in the hybrid expert model.

12. A dual sparse matrix multiplication acceleration device, characterized in that: include: A multiplication element locating module, used for locating a set of compressed matrices and a second sparse matrix matching the first sparse matrix in a global memory of a computing chip according to a sparse matrix multiplication requirement; Wherein, the first sparse matrix includes a plurality of compression units of set size, each compression unit includes at least one valid row and at least one sparse row, each valid row includes at least one structured sparse computing storage unit; each structured sparse computing storage unit is in the same sparse mode; the compression matrix set includes a compression data matrix for storing non-zero data in the first sparse matrix, an index matrix for storing the position of the valid rows in the first sparse matrix in the corresponding compression unit, and a metadata matrix for storing the position of the non-zero data in the corresponding structured sparse computing storage unit; the sparse rows are the matrix rows in which all data are 0, and the sparse mode is the proportion of non-zero data in the structured sparse computing storage unit to all data; A step-by-step transfer module is used to transfer the compressed matrix set and the second sparse matrix from the global memory to the hardware register of the computing chip step by step in the form of data blocks; The multiplication calculation module is used to gradually calculate the multiplication results of the first sparse matrix and the second sparse matrix according to the data loaded in batches in the hardware register through the sparse calculation unit of the calculation chip.

13. An electronic device, characterized in that: The electronic device comprises: at least one computing chip; and A memory communicatively connected to the at least one computing chip; wherein, The memory stores a computer program that can be executed by the at least one computing chip, and the computer program is executed by the at least one computing chip so that the at least one computing chip can execute the multiplication acceleration method of double sparse matrices described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the multiplication acceleration method of double sparse matrices described in any one of claims 1-11 when executed by a computing chip.

15. A computer program product, characterized in that The computer program product comprises a computer program which, when executed by a computing chip, implements the method for accelerating the multiplication of double sparse matrices according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Data processing method and accelerator suitable for sparse neural network calculation array

    CN114970810A

  • FPGA-based graph convolutional neural network sparse matrix multiplication distribution system

    CN115390788A