Arithmetic device, arithmetic method and board card
By introducing a scheduling unit and a cache unit into the computing device, data cache is optimized according to the number of multiplications of matrix elements, the problems of slow and low efficiency of matrix multiplication operations in the prior art are solved, and more efficient computing performance is achieved.
Patent Information
- Application Number
- CN202510608557.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing computing devices have slow operation speed and low efficiency during matrix multiplication, making it difficult to meet the computing needs of some task scenarios.
A computing device is provided, including a scheduling unit, a cache unit and a computing unit. The scheduling unit determines the number of multiplications of each row in the second matrix based on the non-zero elements in the first matrix and indicates it to the cache unit. The cache unit reads and caches the non-zero elements according to the multiplication times, and optimizes the retention and release of data in the cache unit.
By determining and cacheing matrix elements participating in the operation in advance, repetitive data readings are reduced, computing speed and efficiency are improved, and the overall performance of the computing device is improved.
Smart Images

Figure CN120123630A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of data processing, and in particular, to an arithmetic device, an arithmetic method, and a board card. Background Art
[0002] With the development of technologies such as computer vision algorithms, machine learning algorithms, and emerging large artificial intelligence models, large-scale matrix operations are required in many task scenarios. In particular, there is an increasing demand for the operation efficiency during matrix multiplication.
[0003] When the existing arithmetic devices perform matrix multiplication, there are problems of slow operation speed and low operation efficiency, resulting in difficulty for the existing arithmetic devices to meet the operation requirements of some task scenarios. Summary of the Invention
[0004] The embodiments of the present application provide an arithmetic device, an arithmetic method, and a board card, which are used to quickly perform matrix multiplication and improve the operation speed and operation efficiency.
[0005] In a first aspect, the embodiments of the present application provide an arithmetic device, which is applied to the multiplication operation of a first matrix multiplied by a second matrix. The arithmetic device includes a scheduling unit, a caching unit, and at least one arithmetic unit;
[0006] The scheduling unit is configured to determine the number of multiplication times for any row in the second matrix participating in the multiplication operation according to the non-zero elements in the first matrix, and indicate the number of multiplication times for each row in the second matrix to the caching unit;
[0007] The caching unit is configured to read and cache the non-zero elements of each row in the second matrix according to the number of multiplication times for each row;
[0008] The scheduling unit is further configured to input each non-zero element in any row of the first matrix into one of the arithmetic units;
[0009] The arithmetic unit is configured to read the non-zero elements of the corresponding rows in the second matrix participating in the multiplication operation in the caching unit according to each non-zero element in any row of the first matrix, and perform matrix multiplication and combination operations to obtain the corresponding elements in the result matrix.
[0010] In a second aspect, the embodiments of the present application provide an arithmetic method, which is applied to the arithmetic device as described in the first aspect. The arithmetic device includes a scheduling unit, a caching unit, and at least one arithmetic unit. The method is used for the multiplication operation of a first matrix multiplied by a second matrix, and includes:
[0011] Determine the number of multiplication operations for each row in the second matrix participating in the multiplication operation according to the non-zero elements in the first matrix, and indicate the number of multiplication operations for each row in the second matrix to the cache unit;
[0012] Read and cache the non-zero elements of each row in the second matrix according to the number of multiplication operations for each row;
[0013] Input each non-zero element of any row in the first matrix into one of the operation units;
[0014] According to each non-zero element of any row in the first matrix, read the non-zero elements of the corresponding rows in the second matrix participating in the multiplication operation in the cache unit, and perform the multiplication combination operation of the matrices to obtain the corresponding elements in the result matrix.
[0015] In a third aspect, an embodiment of the present application provides a board card, and the board card includes the operation device as described in the first aspect.
[0016] The operation device, operation method, and board card provided by the embodiments of the present application. The operation device can be used for the multiplication operation of multiplying a first matrix by a second matrix. The operation device includes a scheduling unit, a cache unit, and at least one operation unit. During the operation, the scheduling unit can determine in advance the number of multiplication operations for each row in the second matrix participating in the multiplication operation according to the non-zero elements in the first matrix, and can indicate the number of multiplication operations for each row to the cache unit, so that the cache unit can accurately know in advance the number of times the elements in each row of the multiplier matrix participate in the multiplication operation. In this way, when the cache unit reads and caches the non-zero elements of each row in the second matrix, it can accurately decide which elements cached in the cache unit should be retained and which should be deleted as soon as possible according to the number of multiplication operations for each row. This can make the elements of the rows that still need to participate in the operation be retained in the cache unit for as long as possible, avoiding increasing the number of times of repeatedly reading elements due to deleting them in advance, and can also make the elements of the rows that no longer participate in the operation be released as soon as possible, avoiding occupying the cache space for too long. Therefore, making a decision with a high reuse rate of the elements in the cache unit according to the number of multiplication operations can reduce the number of times of repeated data reading and the time consumed by data reading. Based on this, through the close cooperation of the scheduling unit and the cache unit, on-chip data with a high hit rate can be prepared for the operation process of the operation unit as early as possible, so that the operation unit can calculate quickly, thereby improving the overall operation speed and operation efficiency of the operation device. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0018] Figure 1 It is a schematic diagram of matrix multiplication for inner product operation;
[0019] Figure 2 Schematic diagram of matrix multiplication for outer product operation;
[0020] Figure 3 Schematic diagram of matrix multiplication for Gustafson operation;
[0021] Figure 4 One of the structural schematic diagrams of the arithmetic device provided by the embodiment of the present application;
[0022] Figure 5 Another structural schematic diagram of the arithmetic device provided by the embodiment of the present application;
[0023] Figure 6 Schematic diagram of the operation flow of the arithmetic device provided by the embodiment of the present application;
[0024] Figure 7 Schematic diagram of the rearrangement of element positions provided by the embodiment of the present application;
[0025] Figure 8 Schematic diagram of the read-write strategy based on frequency locking provided by the embodiment of the present application;
[0026] Figure 9 Schematic diagram of the multi-level combiner provided by the embodiment of the present application;
[0027] Figure 10 Schematic diagram of the merge tree provided by the embodiment of the present application;
[0028] Figure 11 Schematic diagram of multiple cache slices provided by the embodiment of the present application;
[0029] Figure 12 Schematic diagram of the PFCSR compression method provided by the embodiment of the present application;
[0030] Figure 13 Schematic diagram of the arithmetic method provided by the embodiment of the present application;
[0031] Figure 14 Schematic diagram of a board card provided by the embodiment of the present application.
[0032] Through the above-mentioned drawings, the specific embodiments of the present disclosure have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0033] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0034] In the technical solutions of the embodiments of the present application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0035] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0036] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0037] In the embodiments of the present application, if words such as "first" and "second" are used, it is to distinguish the same items or similar items with basically the same functions and roles. For example, the first electronic device and the second electronic device are only used to distinguish different electronic devices, and their sequence is not limited. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences.
[0038] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0039] To clearly understand the technical solutions of the present application, the prior art will be introduced in detail first.
[0040] Matrices have a wide range of applications in fields such as computer science and engineering. For example, for sparse matrices that include a large number of non-zero elements, general sparse matrix multiplication is widely used in fields such as statistical mathematics, computer vision algorithms, machine learning algorithms, and emerging large models, and the operation efficiency of sparse matrix multiplication has a greater impact on the execution performance of tasks.
[0041] The multiplication between matrices can be simply referred to as matrix multiplication, which includes various operation forms such as inner product operation (Inner Product), outer product operation (Outer Product), and Gustavson operation. The following will introduce matrix multiplication in different operation forms. In Figures 1 to 3 the blank squares represent zero elements in the matrix, and the diagonally filled squares represent non-zero elements in the matrix. Figures 1 to 3 In
[0042] Figure 1 Figure Figure 1 shows a schematic diagram of matrix multiplication for inner product operation. As shown, during the inner product operation, the rows of the first matrix (MatA) and the columns of the second matrix (MatB) perform an inner product to obtain an element at the corresponding position in the result matrix (MatC); traversing and calculating the inner product operation of all rows of MatA and the corresponding columns of MatB can calculate the elements at each position of MatC. After calculating all the elements in MatC, connecting (concat) them according to the position of each element can obtain the complete result matrix MatC. Here, concat can be understood as splicing or combining.
[0043] When performing matrix multiplication based on the inner product operation, it has certain advantages. For example, it has good output locality, and one operation can completely calculate an output point; at the same time, the reuse format of the rows and columns of MatA or MatB is relatively clear, the scheduling is simple, and task splitting and parallelization are easy to implement. When performing matrix multiplication based on the inner product operation, it also has certain disadvantages. For example, since the rows of MatA and the columns of MatB matrix need to be used as inputs, in the case of sparse representation, during the operation, it is also necessary to convert the compressed dense vectors of MatA or MatB into sparse rows with 0s. When the sparsity is low, the utilization rate of the arithmetic unit is extremely low, and the performance is poor.
[0044] Figure 2 Figure Figure 2As shown, during the outer product operation, the columns of the first matrix (MatA) and the rows of the second matrix (MatB) perform the outer product to obtain partial sums in the result matrix (MatC); by traversing and calculating all rows and columns in MatA and MatB, the partial sums at all positions in MatC can be obtained, and by further merging (merging) the partial sums at each position, the complete result matrix MatC can be obtained. Among them, merge can be understood as accumulating the partial sums at each position at their respective positions.
[0045] When performing matrix multiplication based on the outer product operation, there are certain advantages. For example: The outer product operation is based on the multiplication operation between points, has no requirements for the positional relationship between the input and output, and all multiplication operations are effective operations in the sparse representation format, with a relatively high utilization rate of the arithmetic unit; at the same time, parallel operations and input multiplexing can be achieved through the method of row and column splitting. When performing matrix multiplication based on the outer product operation, there are also certain disadvantages. For example: The parallel splitting method of the outer product operation may face a relatively large load imbalance, reducing the overall operation utilization rate; at the same time, the outer product operation requires additional accumulation of the partial sums of MatC. For the accumulation of the partial sums of MatC in the sparse representation, schemes such as hash mapping or heap sorting have relatively high storage space requirements and may not be computable in the case of a large matrix scale; using various comparison sorting methods will cause the partial sum accumulation process to be positively correlated with the output scale and the merging process to be difficult to parallelize, resulting in the performance bottleneck becoming the merging efficiency. The partial sums of MatC need to frequently perform read and write communications with the main memory in the case of a large matrix scale, resulting in an amplification of the task input / output (IO) load and becoming an IO bottleneck.
[0046] Figure 3 It is a schematic diagram of matrix multiplication for Gustavson operation, and the Gustavson operation can also be called row-to-row operation. As Figure 3 shown, an element in any row of the first matrix (MatA) will perform a multiplication operation with an entire row in the second matrix (MatB), and this entire row is the row corresponding to the column of this element in MatA. For example, the element in the 0th row and 0th column of MatA will be multiplied by each element in the 0th row of MatB; for another example, the element in the 0th row and 1st column of MatA will be multiplied by each element in the 1st row of MatB; for another example, the element in the 1st row and 2nd column of MatA will be multiplied by each element in the 2nd row of MatB.
[0047] An element in MatA is multiplied by a whole row of elements in MatB to obtain a whole row of partial sums. The position of each partial sum is the same as the position of the corresponding element in MatB when the partial sum is obtained. These partial sums can be understood as only intermediate results, and they still need to be combined to obtain the elements in the final result matrix (MatC). After combination, all the partial sums of the elements in the same row of MatA will be combined into one row, and the row obtained after combination is in the same row in MatC as the row of the elements in MatA before multiplication. By analogy, traversing each element in MatA and performing multiplication operations with the corresponding rows of MatB, the corresponding partial sums can be obtained. Combining the partial sums of all rows row by row can obtain the elements at each position in MatC. Connecting these elements can obtain the complete result matrix MatC. It should be understood that since multiplying 0 by any number is equal to 0, in the Gustafson operation, only non-zero elements need to be operated on.
[0048] When performing matrix multiplication based on the Gustafson operation, there are certain advantages. For example: similar to the outer product algorithm, all executed multiplication operations are effective operations, and the utilization rate of the computing device is relatively high; compared with the outer product algorithm, by replacing row-column multiplication with row-row multiplication, the entire task operation process is split, and the scale of the partial sums is reduced from the matrix to the row; the locality of the output matrix of this algorithm is also stronger, avoiding the bandwidth pressure and upper cache pressure caused by large-scale data swapping in and out; the task splitting method is also relatively clear. Since the computing granularity is reduced, the operation load balancing strategy is also relatively simple. When performing matrix multiplication based on the Gustafson operation, there are also certain disadvantages. For example: although the data scale is reduced, a combination process is still required to accumulate and combine the partial sums of different rows, and this process still consumes a lot of hardware resources and is difficult to speed up in parallel; when the matrix scale is large or the number of rows is large, the partial sum data still needs to be stored in the main memory for caching, resulting in an increase in the IO bandwidth pressure and affecting the performance; due to the sparse representation of the input data, the locality is poor, and the input data is difficult to reuse.
[0049] When performing the task of matrix multiplication, it can be implemented by a general or special computing device. For example, it can be operated by a general computing device such as a Central Processing Unit (CPU) or a Graphics Processing Unit (GPU), or it can be operated by a matrix multiplication accelerator specifically designed for matrix multiplication.
[0050] When using the existing CPU to process the task of matrix multiplication, since the number of computing units in the CPU is relatively small and the computing power integration is low, the computing concurrency is low, the total amount of computing per unit time is low, and thus the operation efficiency is low when processing the task of matrix multiplication.
[0051] When processing matrix multiplication tasks through existing GPUs, although the GPU has more computing units and a higher degree of computing power integration, it is limited by the GPU's memory capacity and vectorized computing units, and some computing units will be idle, making it difficult to achieve large-scale matrix storage and calculations, and the calculation efficiency is also low.
[0052] Although existing matrix multiplication accelerators are designed or improved specifically according to different forms of matrix multiplication operations, they are limited by the bandwidth bottleneck of data reading. The time-consuming data reading and writing operations will affect the computing speed, resulting in lower computing speed and low computing efficiency.
[0053] For example, an existing matrix multiplication accelerator includes a computing unit, an on-chip cache (such as a cache memory) and a memory (such as a main memory). The computing unit is used to perform multiplication and / or addition operations on the elements involved in the matrix multiplication, the memory is used to store the matrices involved in the matrix multiplication and the intermediate results in the calculation process, and the on-chip cache can be understood as a buffer storage device between the computing unit and the memory, which is used to read and cache the matrix elements that need to participate in the operation from the memory, so that the computing unit can quickly call the elements for calculation.
[0054] The storage capacity of the memory is usually large and can store all the elements of the matrix, but the storage capacity of the on-chip cache is usually small and cannot store all the elements of the matrix at the same time. Therefore, the on-chip cache needs to frequently release the cached data according to the progress of the operation process in order to read the data required for the current operation from the memory. Since the bandwidth between the on-chip cache and the memory is limited, it takes a certain amount of time to read the data. Therefore, if the on-chip cache reads data from the memory more times, it will consume a lot of time, which will cause the calculation speed to be limited by the bandwidth bottleneck, resulting in a decrease in calculation speed and efficiency.
[0055] When two matrices are multiplied, such as MatA multiplied by MatB, MatA can be understood as the matrix to be multiplied and MatB as the multiplier matrix. Taking the on-chip cache as an example, in the operation process of the existing matrix multiplication accelerator, the cache usually decides which row elements currently cached in the cache should continue to be retained on the chip and which should be deleted as soon as possible to free up storage space to cache other elements that need to participate in the operation based on the historical access frequency of each row element of the multiplier matrix.
[0056] This decision-making strategy is based on historical access frequencies, which cannot accurately indicate the subsequent operations during the entire multiplication process. It can be understood that based on the data access situation in the front segment of the operation process, the cache cannot accurately determine which row elements will participate in the operation again in the subsequent operation process, nor can it determine how many times they will participate in the operation, nor can it determine which row elements will no longer participate in the operation.
[0057] Since the on-chip cache cannot accurately perceive and determine in advance the number of times the elements of each row of the matrix participate in the operation, this may cause the on-chip cache to prematurely delete the elements of the rows that will still participate frequently in the operation in the subsequent segment during the operation, or may retain the elements of the rows that no longer participate in the operation on the on-chip for too long. If the elements of the rows that will still participate frequently in the operation in the subsequent segment are prematurely released, then these row elements need to be read from the memory again or multiple times later, which will increase the data reading overhead and reduce the overall operation speed. If the elements of the rows that no longer participate in the operation are retained for a long time, it will meaninglessly occupy the storage space of the on-chip cache for a long time, increasing the storage pressure, and may also, due to the storage of the row elements that no longer participate in the operation and being forced by the storage space pressure, prematurely and wrongly release the row elements that still need to participate in the operation in the subsequent segment, resulting in the need to read these row elements again later and increasing the data reading overhead.
[0058] It can be seen that since the existing matrix multiplication accelerator cannot accurately perceive and determine in advance the number of times the elements of each row of the multiplier matrix participate in the multiplication operation, it will cause the on-chip cache to make suboptimal decisions when retaining and releasing row elements, resulting in a large number of repeated data reading operations, and thus reducing the overall operation speed and efficiency. This situation can also be understood as the overall operation speed and operation efficiency of the multiplication accelerator are reduced due to the low reuse efficiency of on-chip data.
[0059] In view of this, an embodiment of the present application provides an arithmetic device, which includes a scheduling unit, a caching unit, and an arithmetic unit. During arithmetic operations, the scheduling unit can determine in advance the number of multiplication operations in which the elements of each row in the multiplier matrix participate in the multiplication operation, and can indicate the number of multiplication operations for each row to the caching unit, so that the caching unit can more accurately know in advance the number of times the elements of each row in the multiplier matrix participate in the multiplication operation. In this way, when the caching unit reads and caches the row elements of the multiplier matrix, it can more accurately decide whether to keep or discard the on-chip data according to the number of multiplication operations for each row. It can be understood that during arithmetic operations, the row elements that need to participate in the arithmetic operations frequently can be kept in the caching unit for as long as possible, avoiding the increase in the read overhead of reading the element from the outside due to its early release, and the row elements that no longer participate in the arithmetic operations in the later stage can also be released as soon as possible according to the number of multiplication operations, avoiding occupying the caching space. Therefore, the arithmetic device provided by the embodiment of the present application can improve the reuse efficiency of on-chip data during arithmetic operations, reduce the number of times of reading data, reduce the time delay caused by frequent data reading and transfer, speed up the arithmetic speed, and thus improve the arithmetic efficiency.
[0060] The technical solution of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0061] Figure 4 FIG. is a schematic structural diagram of an arithmetic device provided by an embodiment of the present application. The arithmetic device can be applied to the multiplication operation of multiplying a first matrix by a second matrix. For example, Figure 4 as shown, the arithmetic device includes a scheduling unit 401, a caching unit 402, and at least one arithmetic unit 403.
[0062] In the embodiment of the present application, the arithmetic device can be a physical hardware device. For example, it can be an accelerator or an acceleration chip for matrix multiplication operations; or, the arithmetic device can also be a virtual software device, such as a virtual device that realizes matrix multiplication operations through computer programs and instructions; or, the arithmetic device can also be a device combined with a physical part and a virtual part. The embodiment of the present application does not limit this.
[0063] The first matrix and the second matrix are respectively the two matrices participating in the multiplication operation, and these two matrices can be any matrices capable of performing multiplication operations. Any element in the first matrix and the second matrix can be a zero element or a non-zero element. If at least one of the first matrix and the second matrix is a sparse matrix, then the multiplication operation between the first matrix and the second matrix can be understood as the multiplication operation of a sparse matrix. When the first matrix is multiplied by the second matrix, the first matrix can be understood as the matrix to be multiplied, and the second matrix can be understood as the multiplier matrix. The arithmetic device according to the embodiments of the present application can perform the multiplication operation on the first matrix and the second matrix in accordance with arithmetic forms such as Gustavson arithmetic.
[0064] The scheduling unit 401 is configured to determine the number of multiplication times of any row in the second matrix participating in the multiplication operation according to the non-zero elements in the first matrix, and indicate the number of multiplication times of each row in the second matrix to the cache unit.
[0065] Exemplarily, the scheduling unit 401 can be a component in the arithmetic device responsible for managing and allocating computing resources. For example, it can be a Bubble Sort Scheduler (BSS), etc.
[0066] In the embodiments of the present application, the arithmetic device is used to process the multiplication operation of the first matrix multiplied by the second matrix. Therefore, any non-zero element in the first matrix can determine the corresponding element in the second matrix that multiplies with it according to its row pointer, column index, and the matrix multiplication rule. After determining the corresponding element, the number of multiplication times of the row where the corresponding element is located can be determined.
[0067] Wherein, the row pointer can be the row identifier of the row where the element in the matrix is located. For example, the row pointer can be understood as the row number of the row where the element is located; the column index can be the column identifier of the column where the element in the matrix is located. For example, the column index can be understood as the column number of the column where the element is located.
[0068] It should be understood that since the result of multiplying a zero element by any element is zero, the arithmetic device in the embodiments of the present application can only perform corresponding arithmetic processing on the non-zero elements in the matrix, which can improve the arithmetic efficiency.
[0069] Since the corresponding element in the second matrix that multiplies with any non-zero element in the first matrix can be determined according to the non-zero elements in the first matrix, the scheduling unit 401 can count the number of multiplication times of any row in the second matrix participating in the multiplication operation through a statistical computer program or instruction. The number of multiplication times can be understood as the number of times each element in an entire row participates in the operation. For example, if the number of multiplication times of a certain row is a, it means that each element in that row participates in a multiplication operations.
[0070] In a possible implementation, the scheduling unit 401 is specifically configured to: count the number of any column index according to the column index of the non-zero elements in the first matrix; and determine the number of multiplication operations of the rows in the second matrix whose row pointer is the index value of any column index as the number of any column index according to the index value of any column index.
[0071] Exemplarily, since the second matrix is a multiplier matrix, taking the Gustavson operation as an example, if the elements in the i-th column and the j-th column of the x-th row in the first matrix are both non-zero elements, then the element in the i-th column of the x-th row in the first matrix should be multiplied by each element in the i-th row of the second matrix respectively to obtain the partial sum of the elements in the corresponding column of the x-th row in the result matrix; similarly, the element in the j-th column of the x-th row in the first matrix should be multiplied by each element in the j-th row of the second matrix respectively to obtain the partial sum of the elements in the corresponding column of the x-th row in the result matrix.
[0072] Among them, the partial sum can be understood as the intermediate result when gradually accumulating the product for any element during the multiplication operation. For example, by adding all the partial sums of the i-th column of the x-th row in the result matrix during the operation process, the element value of the i-th column of the x-th row in the result matrix can be obtained.
[0073] It can be seen that the number of column indexes of the non-zero elements in the first matrix can represent the number of multiplication operations of the non-zero elements in the second matrix corresponding to the column index. Therefore, according to the index value of any column index, the number of any column index can be determined as the number of multiplication operations of the rows in the second matrix whose row pointer is the index value of any column index.
[0074] For example, assume that the first matrix is a 4×4 matrix, then the row numbers are 0 to 3 respectively. For example, the row pointer of row 0 is 0, the row pointer of row 1 is 1, and so on; similarly, the column numbers of the first matrix are 0 to 3 respectively. For example, the index value of column 0 is 0, the index value of column 1 is 1, and so on. Assume that the second matrix is also a 4×4 matrix, and the row number, column number, row pointer, and index value of the column index are similar to those of the first matrix, and will not be elaborated here.
[0075] According to the column index of the non-zero elements in the first matrix, for example, assume that there are 2 non-zero elements in column 0 of the first matrix, then the number of column 0 is 2; according to the index value of column 0, the number 2 of column 0 can be determined as the number of multiplication operations of the row in the second matrix whose row pointer is 2, that is, the number of multiplications of row 2 in the second matrix is 2. For another example, assume that there are 4 non-zero elements in column 3 of the first matrix, then the number of column 3 is 4. According to the index value of column 3, the number 4 of column 3 can be determined as the number of multiplication operations of the row in the second matrix whose row pointer is 3, that is, the number of multiplications of row 3 in the second matrix is 4.
[0076] It should be understood that in the above examples of the embodiments, only the processing procedures of any one row or two rows are described. In actual applications, the elements of each row in the processing matrix will be traversed in the same or similar manner.
[0077] In the embodiments of the present application, in matrix multiplication, each non-zero element of the first matrix will correspondingly multiply each non-zero element in the row pointed to by the row pointer with the same index value as the column index of the non-zero element. Therefore, by counting the number of column indices of each non-zero element in the first matrix, the number of multiplications for each row in the second matrix can be quickly and accurately determined, improving the efficiency and accuracy of determining the number of multiplications.
[0078] Exemplarily, after determining the number of multiplications for each row in the second matrix, the scheduling unit 401 may send instructions or information to the cache unit 402 to indicate each number of multiplications to the cache unit 402 through the instructions or information.
[0079] The cache unit 402 is configured to read and cache the non-zero elements of each row in the second matrix according to the number of multiplications for each row.
[0080] Exemplarily, the cache unit 402 can be understood as an on-chip cache. The cache unit 402 can interact with the arithmetic unit 403 and input the elements in the second matrix that need to participate in the operation to the arithmetic unit 403. The cache unit 402 can be any one of multiple types of caches, for example, it can be a Distributed Ring Buffer (DRB), etc. The DRB may include several cache slices (such as cache slice), and the cache slice can also be referred to as a cache line (cacheline) or a slice. The cache unit 402 can temporarily store the data and instructions to be accessed to reduce the time delay for the arithmetic unit 403 to access the data.
[0081] The cache unit 402 can interact with the main memory and read the elements in the second matrix from the main memory. The main memory can be understood as a storage unit for long-term data storage, for example, it can be a High Bandwidth Memory (HBM), etc. Before performing the operation on the first matrix and the second matrix, the first matrix and the second matrix can be stored in the main memory first. During the operation process, the elements that need to participate in the operation can be read from the main memory through the cache unit 402.
[0082] After interacting with the scheduling unit 401, the cache unit 402 can receive the number of multiplications for each row in the second matrix indicated by the scheduling unit 401. The cache unit 402 can adopt different read and write strategies according to the number of multiplications for each row to read and cache the non-zero elements of each row in the second matrix.
[0083] Since the larger the number of multiplications, the more times the elements in the corresponding row need to participate in the operation. Therefore, the read-write strategy can be, for example: compare the magnitudes of the numbers of multiplications, and preferentially read and cache the elements in some rows with larger numbers of multiplications in the cache unit 402 for access and invocation by the operation unit 403. During the operation process, if the storage space of the cache unit 402 permits, the elements in the rows with relatively larger numbers of multiplications can be preferentially retained, and the elements in the rows with relatively smaller numbers of multiplications can be preferentially released. In this way, the cache unit 402 can try to retain the elements that need to frequently participate in the operation on the chip, reduce the data reading overhead between the cache unit and the main memory, improve the reuse efficiency of the on-chip data, and thus improve the operation speed and operation efficiency.
[0084] During the operation process, if it is determined that the number of times the operation unit 403 invokes the elements in a certain row has reached the number of multiplications, it indicates that the elements in this row will not participate in the operation subsequently. Then, the cache unit 402 can quickly delete it to release the storage space. According to the numbers of multiplications of each row, the cache unit 402 can also adopt other read-write strategies to read and cache the non-zero elements in each row of the second matrix. The embodiments of the present application do not limit this.
[0085] The scheduling unit 401 is further configured to input each non-zero element in any row of the first matrix into an operation unit 403.
[0086] Exemplarily, during the multiplication operation process, the scheduling unit 401 can also interact with the main memory, read the elements of the first matrix stored in the main memory, and can allocate the non-zero elements in each row of the first matrix to at least one operation unit 403 to perform operations on the elements.
[0087] For example, the scheduling unit 401 can read the non-zero elements in n rows of the first matrix in the main memory, and can allocate these n rows to n operation units 403, and each operation unit 403 is responsible for calculating the non-zero elements in one row. After the operation unit 403 finishes processing one row, the scheduling unit 401 can also allocate the non-zero elements in other rows to enable it to continue the calculation. Until after calculating all the non-zero elements in the first matrix, the scheduling unit 401 can stop inputting elements to the operation unit 403.
[0088] The operation unit 403 is configured to, according to each non-zero element in any row of the first matrix, read the non-zero elements in the corresponding rows of the second matrix participating in the multiplication operation in the cache unit 402, and perform the multiplication and combination operation of the matrices to obtain the corresponding elements in the result matrix.
[0089] Exemplarily, the operation unit 403 can be any computing unit capable of performing multiplication operations and merging operations on the elements of a matrix. If the operation device is a hardware device, the operation unit 403 can be, for example, a unit composed of a multiplication circuit, an addition circuit, a merger, etc. The operation unit 403 can be, for example, a Multiplication Merge Unit (MMU), etc.
[0090] The MMU integrates the multipliers and adders required for matrix multiplication operations. It can achieve concurrent comparison through time-division multiplexing and the idea of frontier comparison, avoiding the repeated data download / storage (load / store) and repeated comparison operations brought by recursive comparison, and can save hardware overhead.
[0091] Taking the Gustafson operation form as an example, the operation unit 403 can, according to the column indices of the non-zero elements in any row of the first matrix input by the scheduling unit 401, call the elements in the row with the row pointer having the same index value as the column index in the cache unit 402, perform multiplication operations on each non-zero element in the first matrix with the corresponding row, and merge the results of the multiplication operations, then the elements at the corresponding positions in the result matrix can be obtained.
[0092] The operation device provided by the embodiments of the present application can be used for the multiplication operation of multiplying a first matrix by a second matrix. The operation device includes a scheduling unit, a cache unit, and at least one operation unit. During the operation, the scheduling unit can, according to the non-zero elements in the first matrix, determine in advance the number of multiplication times for each row in the second matrix participating in the multiplication operation, and can indicate the number of multiplication times for each row to the cache unit, so that the cache unit can accurately know in advance the number of times the elements in each row of the multiplier matrix participate in the multiplication operation. In this way, when the cache unit reads and caches the non-zero elements in each row of the second matrix, it can accurately decide which elements cached in the cache unit should be retained and which should be deleted as soon as possible according to the number of multiplication times for each row. In this way, the elements in the rows that still need to participate in the operation later can be retained in the cache unit for as long as possible, avoiding increasing the number of times of repeatedly reading elements due to their premature deletion, and can also make the elements in the rows that no longer participate in the operation later be released as soon as possible, avoiding their excessive occupation of the cache space for too long. Therefore, making a decision with a high reuse rate for the elements in the cache unit according to the number of multiplication times can reduce the number of times of repeated data reading and the time consumed due to data reading. Based on this, through the close cooperation of the scheduling unit and the cache unit, on-chip data with a high hit rate can be prepared for the operation process of the operation unit as early as possible, so that the operation unit can calculate quickly, thereby improving the overall operation speed and operation efficiency of the operation device.
[0093] Figure 5 This is the second structural schematic diagram of the operation device provided by the embodiments of the present application, asFigure 5 As shown, BSS is a scheduling unit; DRB is a buffer unit, and the DRB includes multiple buffer slices, for example, it can include 16, 32 or more buffer slices. Each buffer slice can cache data such as elements of a matrix; MMU is an arithmetic unit, and the arithmetic device can include multiple MMUs, for example, it can include 16, 32 or more MMUs; HBM is the main memory, which can be understood as the first storage unit and can be used to store the first matrix (MatA) and the second matrix (MatB); the write-back buffer (Write Back Buffer, WBB) can be understood as the second storage unit and is used to store the result matrix (MatC). In some implementations, when the storage space permits, the first storage unit and the second storage unit can also be set to the same storage unit.
[0094] Figure 6 This is a schematic diagram of the operation process of the arithmetic device provided by the embodiment of the present application. As Figure 5 and Figure 6 shown, BSS can read the elements of any row of MatA. For the non-zero elements in a certain row, for example, the non-zero element a im in the i-th row and m-th column of MatA and the non-zero element a in in the i-th row and n-th column of MatA, the row pointer of the row in the second matrix to be multiplied with them can be determined according to the column indexes of the non-zero elements.
[0095] Exemplarily, BSS can determine, according to the column index index a im of a im , through the stream prefetcher that the elements to be multiplied with the non-zero element a im of MatA in MatB are all non-zero elements in the m-th row. Among them, the stream prefetcher can be understood as a unit that can read multiple data streams simultaneously, for example, it can read 64 data streams simultaneously.
[0096] For example, through the stream prefetcher, the element value value a im of a im and the address addr b im in HBM where a non-zero element bm in the m-th row of MatB to be multiplied with a m is stored can be determined. Also for example, through the stream prefetcher, the element value value a in of the non-zero element a in of MatA and the address addr b in in HBM where a non-zero element b n in the n-th row of MatB to be multiplied with a n is stored can also be determined.
[0097] BSS can indicate the address of the non-zero elements of MatB to DRB so that DRB knows where to read data in HBM. DRB will obtain and cache the corresponding non-zero elements in MatB according to the address. BSS can input multiple rows of MatA into multiple MMUs respectively, and each MMU is responsible for calculating the elements of one row, so multiple MMUs can calculate the elements of multiple rows in parallel.
[0098] The MMU can read the non-zero elements of MatB that need to participate in the operation from DRB. Figure 6 As shown, BSS sets the value a of the non-zero element aik in row i and column k in MatA to ik After entering the MMU's multiplier, the MMU can ik Get all non-zero elements of row k of MatB in DRB (using b k ), and these non-zero elements are sequentially added to a ik Multiply them to get the products, which are the partial sums of the corresponding columns of row i in the result matrix.
[0099] For example, value a ik with a non-zero element b in row k of MatB kj The element value b kj Multiply them to get value c ij , the value c ij Represents the element c in row i and column j in MatC ij After all the partial sums of the elements in row i and column j in MatC are calculated, adding all the partial sums can get the element c in row i and column j in MatC. ij The element value of .
[0100] like Figure 3 As shown in the operation process of the Gustafson operation, assuming that a certain MMU is responsible for processing row 0 in MatA, the BSS inputs all non-zero elements of row 0 to the MMU. Assuming that row 0 includes two non-zero elements, namely the element in row 0 and column 0 and the element in row 0 and column 1. The MMU can read all non-zero elements in row 0 and row 1 of MatB in DRB according to the column index of the non-zero element, multiply the elements in row 0 and column 0 of MatA with all non-zero elements in row 0 of MatB in sequence, and multiply the elements in row 0 and column 1 of MatA with all non-zero elements in row 1 of MatB in sequence.
[0101] Assume that the 0th row of MatB contains two non-zero elements, namely the element in the 0th row and 0th column and the element in the 0th row and 3rd column. After multiplying with the element in the 0th row and 0th column of MatA, the partial sum of the 0th row and 0th column and the partial sum of the 0th row and 3rd column of MatC can be obtained. Assume that the 1st row of MatB contains 1 non-zero element, which is the element in the 0th row and 2nd column. After multiplying with the element in the 0th row and 1st column of MatA, the partial sum of the 0th row and 2nd column of MatC can be obtained. It can be seen that if a certain row of MatA contains n non-zero elements, these n non-zero elements will perform multiplication operations with n rows of MatB, and at most n partial sums of rows can be obtained. These n partial sums of rows will be merged through the Merger in the MMU and merged into one row in MatC.
[0102] When merging, the merger will sort the partial sums belonging to the same row in MatC according to the index values of the column indices of each partial sum to merge them into one row of partial sums. During the merging process, if the column indices of some partial sums are the same, indicating that these partial sums are all partial sums of the same element in MatC, the merger will add these partial sums.
[0103] As Figure 6 shown, after the multiplier in the MMU calculates value cij, the MMU can store the column indices of the non-zero elements in the kth row of MatB in the cache. For example, store the column index index b km of the non-zero element b km and the column index index b kn of the non-zero element b kn in the ping buffer and the pang buffer.
[0104] After the multiplier calculates multiple partial sums, these partial sums can be cached for merging. For example, cache the partial sum value c im of the element c im in the i-th row and m-th column of MatC and the partial sum value c in of the element c in in the i-th row and n-th column of MatC, etc. The merger can cooperate with the selector to merge each partial sum.
[0105] For example, if there are several partial sums of the element c ij in the i-th row and j-th column of MatC among the calculated partial sums, when merging, the merger can indicate the column index c kj of the j-th column to the selector, so that the selector can select the element c ij in the i-th row and j-th column of MatC.The respective parts are added. The selector can add these parts to the adder for summation. When summing, the combiner can also indicate a flag bit (such as acc-flag c) to the adder to continuously indicate that the adder sums these partial sums until the flag bit can be reset after the summation is completed. After the adder adds all the partial sums of the element c at the i-th row and j-th column in MatC kj ), the element value of the element c can be obtained. The combiner can store the element value of the element c ij and the column index index c of the element c ij into WBB. ij ij ij
[0106] Exemplarily, in the data stream of data reading and invocation, the access requests of shared resources can be managed by an arbiter (Arbiter, ARB) as in Figure 6 to evenly respond to each data stream.
[0107] Exemplarily, after one row of MatA is input into the MMU, the MMU usually performs multiplication operations on the elements of this row in the order of their appearance, and the order of calculation of each element will not affect the calculation result. Therefore, before allocating the MMU for multiple rows of elements, if the elements of each row are rearranged in position first, and the elements corresponding to the same column index are arranged in the front of their respective rows, then at least some MMUs can read the elements of the row corresponding to the same column index of MatB from the DRB for operation earlier, so that the elements of the row with more multiplication times in the DRB can be called enough times for multiplication as soon as possible, and the DRB can release it as soon as possible, achieving the purpose of using up the reusable elements in the shortest possible time, thereby improving the efficiency of data reuse and the overall operation speed and efficiency of the operation device.
[0108] In one possible implementation, the scheduling unit is specifically configured to: for the non-zero elements in multiple rows of the first matrix, compare the number of column indexes of each non-zero element, and adjust the position of the non-zero element with a larger number of column indexes to the front of the row where the non-zero element is located, so as to preferentially perform multiplication operations on the non-zero element.
[0109] Exemplarily, the non-zero elements in multiple rows can be the non-zero elements in some or all rows of the first matrix. If the reading and processing capabilities of the scheduling unit are relatively high, all the non-zero elements in all rows of the first matrix can be read and processed simultaneously.
[0110] When comparing the number of column indexes of each non-zero element in multiple rows of the first matrix currently processed by the scheduling unit, it can also be understood as comparing the number of non-zero elements corresponding to the column indexes.
[0111] Figure 7 A schematic diagram of the rearrangement of element positions provided by the embodiment of the present application is as follows Figure 7 As shown, the scheduling unit currently processes the non-zero elements in rows 0, 1, 2, and 3 of the first matrix. The numbers in the boxes of each row represent the column index of a non-zero element. For example, the column index of the third non-zero element (counting from right to left) in row 0 is 8, then this non-zero element will perform multiplication operations with all non-zero elements in row 8 of the second matrix; for another example, the column index of the second non-zero element in row 1 is also 8, then this non-zero element will also perform multiplication operations with all non-zero elements in row 8 of the second matrix; for another example, the column index of the first non-zero element in row 3 is also 8, then this non-zero element will also perform multiplication operations with all non-zero elements in row 8 of the second matrix. It can be seen that among the 4 rows from row 0 to row 3 of the first matrix, there are 3 non-zero elements whose column indices are all 8, indicating that the number of column index 8 is 3. Compared with the number of other column indices, the number of column index 8 is larger.
[0112] If the scheduling unit does not rearrange the positions of the non-zero elements in each row and directly inputs each row into the corresponding operation unit, each operation unit will calculate each element in the order of the element positions. Then the operation unit of row 3 will first calculate the non-zero element with column index 8 and will call all non-zero elements in row 8 of the second matrix from the cache unit once; while the operation unit of row 1 will first calculate the non-zero element with column index 3, and after finishing its operation, it will start to calculate the non-zero element with column index 8; and the operation unit of row 0 will first calculate the non-zero element with column index 0, then calculate the non-zero element with column index 5, and then calculate the non-zero element with column index 8. Since these operation units cannot process the elements with the same column index relatively synchronously, they cannot call all non-zero elements in row 8 of the second matrix from the cache unit in a short time, which will cause all non-zero elements in row 8 of the second matrix to remain in the cache unit for a long time to wait for the corresponding operation units of row 0 and row 1 to call during the operation. This will cause the non-zero elements in row 8 to be in a state of waiting to be called on the chip of the cache unit for a long time, which will affect the on-chip data reuse efficiency.
[0113] Therefore, in order to reuse all non-zero elements in row 8 of the second matrix in a short time and improve the on-chip data reuse efficiency, the scheduling unit can, for the non-zero elements in rows 0, 1, and 3 of the first matrix, count and compare the number of non-zero elements corresponding to the same column index according to the number of column indices of each non-zero element, and adjust the positions of the non-zero elements with a large number to the front of the row where the non-zero element is located. Then, multiple operation units can call the reusable elements in each row in the cache unit relatively synchronously, so that the elements in these reusable rows can be called in a short time, improving the on-chip data reuse efficiency.
[0114] Exemplarily, when rearranging the positions of elements, the original position of the element in the row and the number of column indices of the element can be combined for comprehensive rearrangement. For example, the number of elements with column index 8 in rows 0, 1, and 3 is 3; assuming that rows 0, 1, and 3 all include elements with column index 104, the number of column index 104 is also 3. Then the numbers of column index 8 and column index 104 are the same and both are greater than the numbers of other column indices. However, since the original positions of the elements with column index 104 are relatively more backward than the positions of the elements with column index 8, when adjusting and rearranging the positions of the elements, the elements with column index 8 can be preferentially arranged at the beginning of each row, and the elements with column index 104 can be rearranged at the second position of each row. In this way, the position rearrangement of the elements in each row can be more conveniently performed.
[0115] In the embodiments of the present application, by comparing the numbers of column indices of non-zero elements in multiple rows of the first matrix and adjusting the positions of non-zero elements with larger numbers of column indices to the front of the row where the non-zero element is located to preferentially perform multiplication operations on the non-zero element, the cache unit can release reusable data as early as possible, have enough space to read and cache other non-zero elements from the main memory, reduce the data preparation time, improve the data reuse efficiency while also improving the on-chip data hit rate, and further improve the overall operation speed and operation efficiency of the arithmetic device.
[0116] In a possible implementation manner, the scheduling unit may be a BSS. BSS can be understood as an equalization scheduler based on sliding window statistics, which is responsible for evenly and efficiently reusing the matrix multiplication tasks and sending them to each arithmetic unit in units of the rows of MatA. BSS is used to perform frequency-based rearrangement and merging of the index access of MatA rows to MatB rows through sliding window statistics during scheduling and continuous bubble sorting, so as to avoid the problem that the access to MatB rows in the subsequent operation process causes cache miss and reduces the operation efficiency. At the same time, cooperating with DRB can also effectively reduce the off-chip memory access requirements, thereby alleviating the off-chip memory access bottleneck and improving the upper limit of the accelerator performance of the overall arithmetic device.
[0117] The embodiments of the present application take three sliding windows, namely sliding window 1, sliding window 2, and sliding window 3, as examples and combine Figure 7 for illustration. The number of sliding windows can be not limited in specific applications. As Figure 7 shown, the sliding windows 1, 2, and 3 in the dashed boxes can be three sliding windows generated successively at different times. Each sliding window counts the number of column indices of a preset number of non-zero elements in multiple rows within its corresponding time period.
[0118] For example, each sliding window counts a preset number (e.g.,Figure 7 Count the column indices of the 4 (in the middle) non-zero elements, and rearrange the positions of the non-zero elements within the sliding window according to the rule that the larger the number of non-zero elements corresponding to the same column index, the more forward the position. This rule can also be understood as sliding window statistics and continuous bubble sorting.
[0119] As Figure 7 shown, sliding window 1 can count the 4 non-zero elements in each row from row 0 to row 3 (a total of 16 non-zero elements), and obtain the number of corresponding elements in these rows for each column index. During the statistics, a list of each column index can be formed through the column index data stream (stream index). As Figure 7 shown in, the frequency data stream (stream freq) of sliding window 1 counts the number corresponding to each column index during the statistical period of sliding window 1. Sliding window 2 and sliding window 3 are similar. It can be understood that the subsequent sliding windows are generated by sliding backward based on the previous sliding window, and these sliding windows can respectively count the number of the frequency data stream for the 16 non-zero elements within the sliding window.
[0120] As Figure 7 shown, sliding window 2 moves one position backward as a whole based on sliding window 1, and sliding window 3 moves one position backward as a whole based on sliding window 2. The frequency data stream of sliding window 2 and the frequency data stream of sliding window 3 respectively count the number of each column index during the statistical period of their respective sliding windows. It can be understood that since the positions of the elements may be rearranged, index remapping can be implemented during the rearrangement.
[0121] As Figure 7 shown, within each sliding window, the positions of the elements that need to be rearranged can be timely rearranged to form an issue queue. The issue queue can be understood as the new position sorting obtained after the sliding window rearranges the positions of the non-zero elements. After all the sliding windows are completed, the elements of each row can be input to multiple arithmetic units according to the issue queue of the last sliding window. In this way, the arithmetic unit can perform operations on each element in turn according to the finally rearranged elements.
[0122] As Figure 7 shown, in the issue queue corresponding to sliding window 3, for row 0, row 1, and row 3, the element with column index 8 is ranked first; for row 1, row 2, and row 3, the element with column index 23 is ranked among the top, and so on. Compared with before the position rearrangement, elements with column index 8, elements with column index 23, etc. can perform operations approximately synchronously, so that the cache unit can release the called elements as early as possible, reduce the pressure on the storage space, and improve the efficiency of data reuse.
[0123] In this embodiment, multiple sliding windows can be used to slidably rearrange the positions of non-zero elements in multiple rows, which can moderately reduce the computational overhead of the scheduling unit during position rearrangement, enabling the bubble sort algorithm to be applicable to large-scale matrix multiplication tasks and improving the generality of position rearrangement.
[0124] Exemplarily, if the rows input to multiple arithmetic units are in the rearranged order according to the number of column indices, but the cache unit does not read and cache data from the main memory according to the size of the number of column indices, after each arithmetic unit starts to operate, since the elements of each row in the second matrix may not be cached in the cache unit in the rearranged order, it may cause the arithmetic unit to be unable to read the non-zero elements that need to participate in the operation from the cache unit, resulting in a situation of data read miss or low hit rate. Furthermore, it is necessary to wait for the cache unit to read the elements that need to participate in the operation from the main memory, which will increase the time consumption and lead to a decrease in operation efficiency.
[0125] In a possible implementation manner, the scheduling unit is further configured to: perform a descending order sorting on the number of column indices of each non-zero element in the first matrix to obtain the sorting result of each column index, and indicate the sorting result to the cache unit; specifically, the cache unit is configured to: sequentially read and cache the non-zero elements of each row in the second matrix according to the sorting result.
[0126] Exemplarily, the scheduling unit can count the number of column indices of non-zero elements in some or all rows of the first matrix, and perform a descending order sorting on the column indices according to the order from large to small in terms of the number. It can be understood that the sorting of the number of each column index in the first matrix is equivalent to the sorting of the number of row pointers of each row in the second matrix, and can also be understood as the sorting of the number of multiplication times that each row in the second matrix participates in the operation.
[0127] For example, count the number of column indices of non-zero elements in 4 rows from row 0 to row 3 of the first matrix. Among them, the number of column index 8 is 3, the number of column index 5 is 2, the number of column index 23 is 2, the number of column index 2 is 1, and so on; then the sorting result of each column index is 8, 5, 23, 2, etc. Among them, for multiple column indices with the same number, they can be randomly sorted, or sorted in descending order according to the size of the index value of the column index, or sorted in other ways.
[0128] After obtaining the sorting result of each column index, the scheduling unit can indicate the sorting result to the cache unit through data interaction with the cache unit.
[0129] After obtaining the sorting result, the cache unit can sequentially read and cache the non-zero elements of each row corresponding to the row pointer of the column index in the second matrix according to the order of each column index in the sorting result.
[0130] In the embodiment of the present application, by sorting the number of column indexes of non-zero elements in the first matrix in descending order, a descending order result of the multiplication times of each row in the second matrix can be obtained. Reading and caching the non-zero elements of each row in the second matrix in sequence according to this sorting result can cooperate with the rearrangement of the positions of the elements in each row of the first matrix, improve the hit rate of the cache unit during operation, and then can calculate the rows that need to participate in the operation multiple times as soon as possible, improving the operation speed and operation efficiency.
[0131] Exemplarily, in order to improve the on-chip data reuse efficiency of the cache unit, a read-write policy based on a frequency lock can be used to determine which data among the stored data will be preferentially retained and which data will be preferentially replaced.
[0132] In a possible implementation manner, the cache unit is further configured to: determine the multiplication times of any non-zero element in a row of the second matrix as the multiplication times of the any non-zero element; if any arithmetic unit reads a non-zero element from the cache unit once, subtract one from the multiplication times of the read non-zero element; in the case where the cache unit is full and new non-zero elements need to be read and cached, by comparing the multiplication times of the cached non-zero elements with the multiplication times of the new non-zero elements, preferentially cache the non-zero elements with larger multiplication times.
[0133] Exemplarily, during operation, all non-zero elements in a certain row of the second matrix need to be multiplied by a non-zero element in the corresponding first matrix of that row. Therefore, the multiplication times of the row can also be understood as the multiplication times of each non-zero element in that row. After the cache unit obtains the multiplication times of the row, when reading and caching the elements of that row, it can mark the multiplication times of the row as the multiplication times of the elements for each non-zero element in that row. And the cache unit can also gradually decrease the multiplication times one by one based on the number of times each non-zero element is called by the arithmetic unit to obtain the updated multiplication times of the non-zero element, and the updated multiplication times can be understood as the remaining times that the non-zero element will be called later.
[0134] The new non-zero element can be understood as a non-zero element that is not currently cached in the cache unit. In the case where the cache unit is full and new non-zero elements need to be read and cached, in order to preferentially cache the non-zero elements with larger multiplication times, by comparing the multiplication times of the cached non-zero elements with the multiplication times of the new non-zero elements, preferentially cache the non-zero elements with larger multiplication times. Since the non-zero elements with relatively larger multiplication times are called more frequently and have a higher probability of being called, the comparison based on the multiplication times can improve the data reuse efficiency.
[0135] Figure 8 Schematic diagram of the read-write policy based on a frequency lock provided for the embodiment of the present application, as Figure 8As shown, the elements of each row of the second matrix MatB are stored in the HBM. MatB includes 5 rows, namely row 0, row 1, row 2, row 3, and row 4, which can be respectively identified as R0 to R4 for easy description. Each non-zero element included in any row can be represented by a different serial number. For example Figure 8 within the box of Figure 8 , R0-0 represents the first non-zero element of row 0, and R0-1 represents the second non-zero element of row 0; for another example, R1-2 represents the third non-zero element of row 1, and so on, which will not be elaborated here.
[0136] Assume that the multiplication times corresponding to each row from row 0 to row 4 are 2, 1, 2, 3, and 2 respectively. According to the read and write strategy of the cache unit to read and cache each row, the elements that need to participate in the operation can be read, cached, and called respectively through the following steps S1-S9 according to the operation requirements of the operation unit. It should be understood that when an element is read from the main memory to the cache unit, it will be called by the operation unit for operation once. Therefore, when reading and caching non-zero elements, the multiplication times of the non-zero elements should be reduced by one on the basis of the multiplication times at the time of reading when caching.
[0137] As Figure 8 shown, the step of S1 is to read all non-zero elements of row 0. When reading, since there is no cached data in the cache unit before, the storage space of the cache unit is relatively sufficient, and all non-zero elements of row 0 can be cached. And since the original multiplication times of the non-zero elements of row 0 are 2 and they are called by the operation unit once when being read, after subtracting 1, the multiplication times of the non-zero elements of row 0 are decreased to 1, as Figure 8 the number 1 in the box below R0-0 in Figure 8 represents the decreased multiplication times.
[0138] The step of S2 is to read all non-zero elements of row 1. Since the current remaining storage space of the cache unit can meet the storage space requirements of all non-zero elements of row 1, it is not necessary to delete any non-zero element of the cached row 0.
[0139] In the step of S3, the operation unit calls all non-zero elements of row 0 from the cache unit to participate in the operation, then the multiplication times of the non-zero elements of row 0 are decreased from 1 to 0 after the elements are read. At this time, the storage space of the cache unit is full, and the multiplication times of all non-zero elements are also reduced to 0. In the next step, these non-zero elements can be deleted sequentially or randomly to make room for storing the elements that need to participate in the subsequent calculation.
[0140] In the step of S4, the operation unit needs to call all non-zero elements of row 2 to participate in the calculation. At this time, the cache unit will read and cache these elements according to the addresses of the non-zero elements of row 2 in the main memory. As Figure 8As shown, the 2 rows include 2 non-zero elements. The initial number of multiplications for these two non-zero elements is 2. When they are read, they are called once by the arithmetic unit. After being cached in the cache unit, the number of multiplications for these two non-zero elements decreases to 1. Also, in order to cache the two non-zero elements of the 2 rows, the two non-zero elements of the 0 row are deleted at this time to cache the two non-zero elements of the 2 rows.
[0141] Similar to step S4, in step S5, multiple elements with a multiplication count of 0 are deleted to cache each non-zero element of the 3 rows. In step S5, the arithmetic unit calls each non-zero operation of the 2 rows for calculation. Therefore, the multiplication counts of R2-0 and R2-1 decrease to 0.
[0142] In step S6, it is necessary to read and cache each non-zero element of the 4 rows. Since there are many non-zero elements in the 4 rows, the storage space of the current cache unit cannot meet the requirement of caching all non-zero elements of the 4 rows in the cache unit. Therefore, it is necessary to decide which elements to cache in the cache unit by comparing the multiplication counts of the cached non-zero elements with those of the new non-zero elements.
[0143] As Figure 8 shown, the multiplication counts of the 4 non-zero elements R2-0, R2-1, R1-5, and R1-6 are 0, which is less than the multiplication count of 1 for each non-zero element in the 4 rows. Therefore, the 4 non-zero elements R2-0, R2-1, R1-5, and R1-6 can be deleted, and each non-zero element of the 4 rows is preferentially cached. Also, since the multiplication count of 1 for each non-zero element in the 4 rows is less than the multiplication count of 2 for each non-zero element in the cached 3 rows, therefore, each non-zero element of the 3 rows is preferentially retained, and only some non-zero elements of the 4 rows are cached. The non-zero elements of the 4 rows that cannot be cached can be not cached and can be read again from the main memory when they need to participate in the calculation later. As Figure 8 shown in, R4-0, R4-1, R4-2, and R4-3 are replaced with R2-0, R2-1, R1-5, and R1-6, and R4-4 that cannot be cached is not cached.
[0144] Steps S7 and S8 are similar. In these two steps, the arithmetic unit calls all non-zero elements of the 3 rows from the cache unit to participate in the operation once. Then, the multiplication counts of each non-zero element of the 3 rows decrease from 2 to 0 after the elements are called 2 times. At this time, the multiplication counts of each non-zero element of the 3 rows are 0, so they can be replaced by new non-zero elements that can be read in the next step.
[0145] In step S9, the arithmetic unit needs to call the non-zero elements of each of the 4 rows. Since R4-0, R4-1, R4-2, and R4-3 have been cached in the cache unit, they can be directly called, and the cache unit can call R4-4 from the main memory again. Since the number of multiplications of R4-4 also reduces to 0 after this call, the cache unit may not cache it in the cache unit, or it can also replace a stored element with a multiplication count of 0 and cache R4-4 in the cache unit. Thus, the operation on rows 0 to 4 of the second matrix is completed.
[0146] As can be seen from the above example, through the read-write strategy based on frequency locking, the stored non-zero elements will not be directly replaced before being used up. Instead, by comparing the magnitudes of the multiplication counts, the non-zero elements with larger multiplication counts are preferentially retained, thereby increasing the residence time and reuse times of the non-zero elements on the chip, and further improving the hit rate of reading elements during operation.
[0147] In the embodiment of the present application, based on the read-write strategy of frequency locking, preferentially caching non-zero elements with larger multiplication counts can reduce the premature release of non-zero elements that need to participate in operations frequently, thereby reducing the read overhead between the cache unit and the external memory, and improving the efficiency of data reuse.
[0148] Exemplarily, during the operation process, the combiner in the arithmetic unit is also a factor affecting the operation speed. If the efficiency of the combiner in combining data is higher, the calculation speed of the arithmetic unit can be accelerated.
[0149] In a possible implementation manner, any arithmetic unit includes at least two combiners, and any combiner is used to combine the partial sums of any element in the result matrix, where the partial sum is the intermediate result when the matrix gradually accumulates products for any element during the multiplication operation.
[0150] Exemplarily, during the calculation process of the arithmetic unit, the product a·b obtained by multiplying a non-zero element a of the first matrix by the corresponding non-zero element b of the second matrix is a partial sum. The row pointer of this partial sum is the same as the row pointer of the non-zero element a, and the column index of this partial sum is the same as the column index of the non-zero element b. Then, the partial sum a·b can be understood as a part of the element at the row pointer of a and the column index of b in the result matrix. Adding all the partial sums at the row pointer of a and the column index of b in the result matrix gives the element value at this position.
[0151] Exemplarily, any arithmetic unit may include two or more combiners, and each combiner may be the same combiner or different combiners. For example, an arithmetic unit may include combiners with various different combining capabilities, where the combining capability may be reflected by the number of rows of partial sums that the combiner combines. For example, the combining capability of a combiner that can simultaneously combine 4 rows of partial sums is greater than that of a combiner that can only simultaneously combine 2 rows of partial sums. The combiners for multi-way concurrent sorting can increase the hardware resource overhead in a linear relationship to increase the concurrent combining throughput. The specific implementation of the partial sums and the combiners may also refer to the descriptions in other embodiments of this application.
[0152] In the embodiments of this application, through multiple combiners in the arithmetic unit, the processing of intermediate results can be accelerated in a parallel or multi-level combining manner, the speed and efficiency of combining partial sums can be improved, and thus the overall arithmetic efficiency of the arithmetic device can be improved.
[0153] In a possible implementation, the arithmetic unit includes two combiners. One of the two combiners is a first-level combiner, and the other combiner is a second-level combiner. The first-level combiner is used to serially combine m groups of partial sums to obtain m first-level combined results. Any one of the m groups of partial sums includes n rows of partial sums, and the n rows respectively correspond to n rows in the second matrix, where m and n are both positive integers. The second-level combiner is used to combine the m first-level combined results to obtain a second-level combined result, and the second-level combined result is used to obtain the elements of one row in the result matrix.
[0154] Exemplarily, in the arithmetic unit of this embodiment, two combiners are included, and the two combiners can be configured in a multi-level combining manner to perform multi-level combining processing on multiple rows of partial sums.
[0155] Figure 9 FIG. is a schematic diagram of the multi-level combiner provided for the embodiments of this application, as Figure 9 shown, L1 Merger is the first-level combiner, and L2 Merger is the second-level combiner. The first-level combiner can serially process m groups of partial sums, and each group of partial sums may include n rows of partial sums. The second-level combiner is used to perform secondary combination on the results combined by the first-level combiner.
[0156] For ease of understanding and description, taking m and n both being 4 and combining with Figure 9 for illustration, as Figure 9 shown, each small square represents 1 partial sum. The 16 rows of partial sums are evenly divided into 4 groups, and each group of partial sums includes 4 rows of partial sums. L1Merger can serially process these 4 groups of partial sums in 4 time slots of T0, T1, T2, and T3 by using time-division multiplexing.
[0157] For example, during the time period of T0, the first group of partial sums is merged with the input L1 Merger to output the first first-level merge result; during the time period of T1, the second group of partial sums is merged with the input L1 Merger to output the second first-level merge result; and so on. After the time periods of T2 and T3, a total of 4 first-level merge results can be obtained. These 4 first-level merge results still need to be further merged into one line. Therefore, these 4 first-level merge results can be input into the L2 Merger for further merging to obtain a second-level merge result, which is a line of partial sums, and this line of partial sums can be used to obtain the elements of one line in the result matrix.
[0158] It should be understood that the start time when the second-level merger starts to merge each first-level merge result can be any time after the first-level merger has merged at least one merge result for the last group of partial sums. It can be understood that if there are m first-level merge results participating in the second-level merge, the second-level merger can start to merge when data has been generated in the m-th first-level merge result, without being limited to starting to merge only after the m-th first-level merge result has been completely generated. In this way, the first-level merger and the second-level merger can be merged simultaneously, shortening the time required for merging and improving the merging efficiency.
[0159] In the embodiment of the present application, the arithmetic unit includes two mergers, and the two mergers have different levels. Among them, the first-level merger can initially merge multiple lines of partial sums in a time-division multiplexing manner, and the merged result is handed over to the second-level merger for further merging. In this way, multiple lines of partial sums can be merged in batches with limited hardware resource overhead by the two mergers. Compared with the hardware resource overhead of deploying more mergers, this method can effectively improve the merging efficiency of the mergers while taking into account a relatively small hardware resource overhead, and thus can improve the arithmetic efficiency of the arithmetic device.
[0160] Exemplarily, if the second-level merge result output by the L2 Merger is already the result of having merged all the partial sums of a certain row in the result matrix, then this second-level merge result corresponds to the element values of each element in that row of the result matrix.
[0161] If the second-level merge result output by the L2 Merger is not the result of having merged all the partial sums of a certain row in the result matrix, then this second-level merge result is still an intermediate result, and the element values of the elements in that row of the result matrix can only be obtained after it is merged with the other partial sums of that row. For this situation, considering using the idle merger to further merge multiple second-level merge results, these second-level merge results can be merged as soon as possible to obtain the element values of each element in that row of the result matrix as soon as possible.
[0162] In a possible implementation, the arithmetic device includes a plurality of arithmetic units. When the second matrix includes k rows, where k is a positive integer and the product of m and n is less than k, the scheduling unit is further configured to: call a merger in an idle state among the plurality of arithmetic units to merge partial sums of the plurality of secondary merge results to obtain at least one multi-level merge result, and the multi-level merge result is used to obtain the elements of a row in the result matrix.
[0163] Exemplarily, the product of m and n can be understood as the total number of rows of partial sums that a primary merger in an arithmetic unit currently needs to merge. If the total number of rows is less than the number of rows k of the second matrix, it indicates that the merging ability of the primary merger is insufficient to merge the partial sums corresponding to multiple rows in a certain row of the result matrix within one merging round.
[0164] In this case, by calling a merger in an idle state among the plurality of arithmetic units, the scheduling unit can borrow the mergers that are not in a working state in these arithmetic units to merge the obtained secondary merge results, thereby forming a merger tree structure of multiple mergers, which can improve the merging efficiency.
[0165] Figure 10 The following is a schematic diagram of the merger tree provided by the embodiments of the present application, as Figure 9 and Figure 10 shown. Assume that the number of rows of partial sums that the primary merger of the first arithmetic unit (MMU1) needs to merge is large, and the primary merger can only process 16 rows of partial sums in the first merging round (4 time periods from T0 to T3) and obtain the first secondary merge result. In the second merging round, the primary merger can process another 16 rows of partial sums in the time periods from T4 to T7 and obtain the second secondary merge result. And so on, in the time periods from T8 to T11, it processes another 16 rows of partial sums to obtain the third secondary merge result, and in the time periods from T12 to T15, it processes another 16 rows of partial sums to obtain the fourth secondary merge result.
[0166] The scheduling unit calls a merger in an idle state among the plurality of arithmetic units, for example, calls the merger of the second arithmetic unit (MMU2), then can use the L1 Merger of MMU2 to merge the above 4 secondary merge results, and can obtain a multi-level merge result, which can be understood as a tertiary merge result. After accumulating 4 tertiary merge results, the 4 tertiary merge results can be merged by the L2 Merger of MMU2 to obtain a quaternary merge result. Based on this, by using the mergers in an idle state to participate in the merging of partial sums, the merging efficiency can be effectively improved.
[0167] It should be understood that when the MMU2 combines the second-level combination results, the MMU1 can continue to process the partial sums of other groups in the subsequent period of T15, and the MMU2 can also process other second-level combination results in the period of Tn. Therefore, the combination tree formed by the MMU1 and the MMU2 or the combiners in more computing units can be understood as a topological structure of parallel combination of multiple combiners.
[0168] It can be understood that in the task fields of data processing, image processing, model learning, etc., any matrix participating in the operation may be a large-scale matrix, which may include tens of thousands, millions, or even tens of millions of rows of elements. Therefore, the numbers in the above embodiments are only set for the convenience of illustration. In practical applications, whether it is a first-level combiner or a second-level combiner, the number of rows combined simultaneously is not limited to 4 rows, and a combination round is not limited to processing 4 groups of partial sums. In addition, the computing units or combiners forming the combination tree may not be limited to 2, 3, or 4, and may be a combination tree composed of more combiners to meet the computing requirements of large-scale matrices.
[0169] In the embodiments of the present application, by calling the combiners in the idle state, the combiners in other computing units can be used to combine the second-level combination results, and then a combination tree in a tree structure can be formed, which can process more rows of partial sums in parallel, thereby improving the combination speed and efficiency, and achieving the purpose of improving the computing speed and computing efficiency of the computing device.
[0170] The computing device provided in the embodiments of the present application can reduce the total off-chip IO volume and relieve the bandwidth pressure by increasing the number of rows processed simultaneously by the combiners in the computing units and reducing the number of iterations required to combine all partial sum rows.
[0171] Existing computing devices or accelerators have the defects of redundant partial sum storage in the matrix multiplication calculation task and the defects of hardware power consumption and area constraints. However, the computing device of the embodiments of the present application can improve the overall combination ability of the computing units through multi-level combiners or combination trees, can combine multiple rows of partial sums simultaneously, and thus can reduce the number of times of redundant storage of partial sums during combination, overcoming the defect of redundant partial sum storage. The computing device of the embodiments of the present application can effectively reduce the hardware resource overhead of the combiners through time-division multiplexing combiners or combination trees composed of multiple combiners, overcome the defects of hardware power consumption and area constraints, and can also increase the maximum number of rows combined by the combiners. For example, it can process data of 65,536 rows or even more rows, solving the problem of iterative memory access calculation overhead caused by the inability to completely combine on-chip.
[0172] In a possible implementation, the cache unit includes p groups of buffer areas, each group of buffer areas includes q buffer slices, and any buffer slice is used to cache non-zero elements in the second matrix. Any buffer slice corresponds to an arithmetic unit, where p and q are both positive integers. The scheduling unit is specifically configured to: when it is necessary to call the mergers in the idle state among multiple arithmetic units, preferentially call the arithmetic units corresponding to the buffer slices in the same group of buffer areas.
[0173] Exemplarily, p groups of buffer areas can be divided in the cache unit, and each buffer area can include q buffer slices, where p and q are both positive integers. A buffer slice can correspond to an arithmetic unit, and the arithmetic unit can read the non-zero elements to be involved in the operation through the buffer slice corresponding to it.
[0174] Figure 11 The figure is a schematic diagram of multiple buffer slices provided by the embodiments of the present application. As Figure 11 shown, in the DRB, the boxes numbered 0-3 respectively represent the buffer slices numbered 0 to 3, and these 4 buffer slices form the first group of buffer areas; the boxes numbered 4-7 respectively represent the buffer slices numbered 4 to 7, and these 4 buffer slices form the second group of buffer areas; the boxes numbered 8-11 respectively represent the buffer slices numbered 8 to 11, and these 4 buffer slices form the third group of buffer areas; the boxes numbered 12-15 respectively represent the buffer slices numbered 12 to 15, and these 4 buffer slices form the fourth group of buffer areas. Figure 11 The solid arrows in the figure represent multiple buffer slices in the same buffer area, and the multiple buffer slices in the same buffer area can quickly call data with each other.
[0175] Since the buffer slices in the same group of buffer areas are adjacent to each other, the arithmetic units corresponding to these buffer slices can call parameters such as the element values and column indexes of non-zero elements with low latency. Therefore, when it is necessary to call the mergers in the idle state among multiple arithmetic units, the scheduling unit can reduce the time cost of data call and improve the merging efficiency by preferentially calling the arithmetic units corresponding to the buffer slices in the same group of buffer areas.
[0176] In the embodiments of the present application, a buffer area refers to multiple buffer slices divided into a group. When forming a merging tree, if the mergers corresponding to the arithmetic units of the buffer slices in the same group of buffer areas are preferentially called, the operation requirements of matrix multiplication can be met at the cost of lower hardware resource overhead, achieving the purpose of improving operation efficiency.
[0177] To improve the overall data reading and writing efficiency of the cache unit, data paths can be established among the buffer areas.
[0178] In a possible implementation, between p groups of buffer areas, a data path is set between any two adjacent groups of buffer areas, and the data path is used to transmit non-zero elements cached in the buffer slices.
[0179] Exemplarily, the data path can be a communication channel set for transferring data between buffer slices. If a data path is set between any two adjacent groups of buffer areas among p groups of buffer areas, then the non-zero elements cached in the buffer slices can be transmitted through the data path. In this way, for the arithmetic units corresponding to these buffer slices, when they need to call some non-zero elements, they can call the non-zero elements cached on the buffer slices that do not correspond to them through the data path, improving the hit rate of data reading.
[0180] As Figure 11 shown by the dotted arrows in the figure, a data path can be established among buffer slices No. 0, No. 4, No. 8, and No. 12; a data path can be established among buffer slices No. 1, No. 5, No. 9, and No. 13; a data path can be established among buffer slices No. 2, No. 6, No. 10, and No. 14; a data path can be established among buffer slices No. 3, No. 7, No. 11, and No. 15. In this way, a ring-distributed cache structure can be formed among the 4 buffer areas of the first to the fourth groups. The whole cache is sliced and interconnected through a ring path, which can replace the cache topology structure of the crossbar structure, effectively reducing the communication delay during local access and improving the pipeline performance.
[0181] In the embodiments of the present application, by setting a data path between buffer areas, the non-zero elements cached can be pushed through the data path, so that all buffer areas can form a ring structure. This structure can replace the complex structure of the crossbar, reduce the local access latency, and improve the data transfer performance of the buffer unit.
[0182] It can be seen from the above embodiments that the buffer unit with a ring-separated topology structure can effectively reduce the communication delay between buffer slices on the chip and save the crossbar overhead of loading data between the cache and multiple arithmetic units.
[0183] Exemplarily, during the operation process, there will always be data communication between the buffer unit and the main memory. Even if the number of invalid data repeated readings is reduced through some algorithms or strategies, the overall operation efficiency of the computing device will still be affected by the bandwidth bottleneck between the buffer unit and the main memory. Therefore, in order to further reduce the impact of the bandwidth bottleneck on the overall operation efficiency, the embodiments of the present application can perform corresponding compression according to the type of the second matrix to reduce the amount of data transmitted between the buffer unit and the main memory, and achieve the purpose of improving the operation speed by reducing the amount and increasing the speed.
[0184] In a possible implementation, the arithmetic device includes a first storage unit and a second storage unit. The first storage unit is used to store a second matrix, and the second storage unit is used to store a result matrix; when the second matrix is a sparse matrix, the first storage unit is used to store a compressed second matrix obtained by compressing the sparse matrix, and the cache unit reads the non-zero elements of each row in the compressed second matrix from the first storage unit.
[0185] Exemplarily, if the first storage unit is used to store the elements in the second matrix, the first storage unit can also be understood as the main memory in the above embodiments, such as HBM. The second storage unit is used to store the result matrix, so the second storage unit can also be understood as the write-back buffer WBB in the above embodiments.
[0186] A dilution matrix can be understood as a matrix in which many elements are 0. Since 0 multiplied by any number is equal to 0, therefore, zero element compression can be performed on a sparse matrix, and the calculation result will not be affected after compression, and the data volume of the matrix can also be reduced.
[0187] For the case where the second matrix is a sparse matrix, when storing the second matrix in the first storage unit, the sparse matrix can be compressed first to obtain a compressed second matrix, and then stored in the first storage unit. In this way, not only can the storage space of the first storage unit occupied by the second matrix be reduced, but also the data volume of reading and writing data between the cache unit and the first storage unit can be reduced during the operation.
[0188] Exemplarily, when compressing the second matrix into a sparse matrix, any method capable of compressing and storing the sparse matrix can be used for compression. For example, it can be compressed and stored in the form of a coordinate list (COO), or in the form of a compressed sparse row (CSR), or in the form of a compressed sparse column (CSC). The embodiments of the present application do not limit this.
[0189] In the embodiments of the present application, the first storage unit can store the compressed sparse matrix, which can reduce the space occupancy rate of the second matrix on the first storage unit, save the storage space of the first storage unit, and can reduce the data volume of data transmission between the cache unit and the first storage unit, improve the data transmission efficiency, and is beneficial for the arithmetic unit to quickly read the non-zero elements that need to participate in the operation, and improve the overall operation speed of the arithmetic device.
[0190] Since the existing sparse matrix compression method has a low compression rate, it is impossible to achieve matrix compression with a high compression rate. Therefore, the embodiment of the present application provides an improved CSR compression method for compressing sparse rows CSR, namely, a prefix compressed sparse row (Prefix Compressed Sparse Row, PFCSR) compression method, to further compress the second matrix and reduce the amount of data when reading the second matrix.
[0191] Exemplarily, the compression method of the PFCSR can be used for the first matrix and / or the second matrix, and the first matrix and / or the second matrix can be sparsely represented to obtain the corresponding matrix after compression. Figure 12 The following description is made by taking the compression of the second matrix as an example.
[0192] In a possible implementation, the compressed second matrix is obtained by compressing the sparse matrix in the following manner: compressing the second matrix by compressing the sparse row CSR to obtain a preliminarily compressed second matrix; for the column index of each non-zero element in any row of the preliminarily compressed second matrix, the column index is divided according to the bits of a preset length to obtain at least two levels of sub-indexes, and any two adjacent levels of sub-indexes have a superior-subordinate relationship; for any level of sub-index in the at least two levels of sub-indexes, the sub-indexes with the same sub-index of the upper level and the same sub-index of the current level are merged into one sub-index, and the correspondence between the merged sub-index and the sub-index of the same column index in the sub-index of the lower level is marked; the column indexes in the preliminarily compressed second matrix are replaced by the merged sub-indexes and the marked correspondence to obtain a compressed second matrix.
[0193] Exemplarily, compressed sparse row CSR is a method of storing a sparse matrix through a value array, a column index array, and a row pointer array. Among them, the value array stores the values of all non-zero elements, the column index array stores the column index of each non-zero element, the row pointer array stores the index of the first non-zero element of each row in the value array, and the length of the row pointer array is equal to the number of rows of the matrix plus one. Usually, the compression rate of the row pointer array in the compressed sparse row CSR is higher than the compression rate of the column index array, and there is still a certain amount of compressible space in the column index array. The PFCSR compression method is to further compress each column index in the column index array obtained after compressing the sparse row CSR.
[0194] The preset length of bits can be one of multiple preset lengths, for example, it can be a preset length of 4 bits, 8 bits or 16 bits. Figure 12 The PFCSR compression method is described with examples.
[0195] Figure 12A schematic diagram of the PFCSR compression method provided in the embodiment of the present application is shown in FIG. Figure 12 As shown in the figure, assuming that each column index in the column index array is a hexadecimal number, the length of each column index is 32 bits. For example, a row includes 6 non-zero elements, and the column indexes of these non-zero elements are "00000000", "00000005", "000001EF", "000101CF", "000101FE" and "001039E3". It can be understood that each column index is a 32-bit original column index. It can be seen that the 6 column indexes occupy a length of 32 bits × 6 = 192 bits.
[0196] For any column index, at least two levels of sub-indexes can be obtained by splitting it according to the preset length. For example, the 32-bit orignal column index "001039E3" can be split according to the preset length of 8 bits to obtain 4 levels of sub-indexes, namely the first-level sub-index "E3", the second-level sub-index "39", the third-level sub-index "10" and the fourth-level sub-index "00". Splitting the column index according to the preset length can also be understood as prefix decomposition. The first-level sub-index can also be understood as an 8-bit cut-down index; the second-level sub-index can also be understood as an 8-bit L2 prefix; the third-level sub-index can also be understood as an 8-bit L1 prefix; and the fourth-level sub-index can also be understood as an 8-bit row id. Since the splitting is performed according to the bits of the column index value, the resulting sub-indexes have a hierarchical relationship, which can be understood as the relationship between high bits and low bits. The relationship between sub-indexes at all levels can help restore the compressed data to the original column index representation when the data is decompressed.
[0197] For any level of sub-indexes in at least two levels of sub-indexes, merge the sub-indexes with the same upper level sub-index and the same level sub-index into one sub-index, and mark the corresponding relationship between the merged sub-index and the sub-indexes with the same column index in the lower level sub-indexes. This step can be understood as prefix compression.
[0198] like Figure 12 As shown, for example, 2 "00"s in the second-level sub-index are merged into 1 "00", and 2 "01"s are merged into 1 "01"; for another example, 3 "00"s in the third-level sub-index are merged into 1 "00", and 2 "01"s are merged into 1 "01".
[0199] In the case where there is no upper - level sub - index for any level of sub - index, the same sub - indices in that any level of sub - index are merged into one sub - index. For example, if there is no upper - level sub - index for the fourth - level sub - index, then without considering whether its upper - level sub - indices are the same, the 6 "00"s in the fourth - level sub - index can be directly merged into 1 "00".
[0200] The corresponding relationship between the merged sub - index at each level and the sub - indices in the lower - level that are the same - column indices can be marked in each level of sub - index. For example, the corresponding relationship between the upper - level sub - index and the lower - level sub - index can be indicated by a base - address offset. This base - address offset can be understood as one upper - level sub - index after merging corresponding to several sub - indices in the lower - level.
[0201] Such as Figure 12 In it, the base - address offset "02" is added to the "00" of the second - level sub - index, indicating that the second - level sub - index "00" corresponds to 2 sub - indices in the first - level sub - index. That is, the second - level sub - index "00" is the upper - level sub - index of the 2 sub - indices "00" and "05" in the first - level sub - index. During decompression, "00" will be added to the front of "00" and "05" respectively to restore the column - index representation before compression. Another example, the base - address offset "02" is added to the "01" of the second - level sub - index, indicating that the second - level sub - index "01" corresponds to 2 sub - indices in the first - level sub - index. That is, the second - level sub - index "01" is the upper - level sub - index of the 2 sub - indices "CF" and "FE" in the first - level sub - index. During decompression, "01" will be added to the front of "CF" and "FE" respectively to restore the column - index representation before compression, and the rest will not be elaborated.
[0202] From Figure 12 The above example can be seen that after the 6 original column indices are split and merged after compression, the sub - indices and the base - address offsets are respectively "00", "05", "EF", "CF", "FE", "E3", "02", "01", "02", "01", "00", "01", "01", "39", "02", "01", "01", "00", "01", "10", "03" and "00", a total of 22. Each sub - index or base - address offset occupies 8 - bit length. Then, after compression, it occupies a total length of 8 bit×22 = 176 bit, which is 16 bit less than the 192 - bit length before compression.
[0203] Exemplarily, for the second matrix, after merging and compressing the sub - indices at each level in the second matrix, replacing each column index in the preliminarily compressed second matrix with the merged sub - indices and the marked corresponding relationships, the compressed second matrix can be obtained.
[0204] Existing dilution representation methods for sparse matrices, such as BCSR or TileSpGEMM representation methods, transform the CSR of a large matrix into a blocked CSR, thereby reducing the scale of simultaneous calculations. However, after blocking, the matrix still needs to perform loop operations at the block granularity, and there are still repeated IO operations during the merging process, which cannot improve the operation efficiency. Or, the dilution representation method of C2SC is to continue to perform secondary sparse representation on the row pointers on the basis of the CSC format, thereby compressing the storage space, but does not decompose the column index. At the same time, this method increases the complexity of operation scheduling and decoding and is only applicable to specific hardware structures, which is not flexible and efficient enough.
[0205] In order to further improve the storage performance of sparse matrix multiplication and provide a data format suitable for hardware acceleration, an embodiment of the present application proposes a new sparse matrix representation format PFCSR. Based on the basic compressed sparse row CSR, PFCSR further performs prefix decomposition of the column index vector in the original CSR, splits the column index into at least two levels of sub-indexes, and at the same time, for the same sub-indexes in the same-level sub-indexes, further performs format compression representation through merging, and finally represents the compressed format of PFCSR. By testing the compression ratio of the compressed format of PFCSR on the general sparse matrix multiplication benchmark SuiteSparseMatrix Collection, a compression of an average of 20% of the original matrix data volume is achieved, indicating that the compressed format of PFCSR can further improve the matrix compression ratio compared with the general compressed sparse row CSR compression format.
[0206] The sparse representation method of PFCSR provided by the embodiment of the present application is to divide the column index into multiple levels of sub-indexes based on a preset length of bit positions, merge the same sub-indexes and mark the corresponding relationships of the sub-indexes originally belonging to the same column index, so as to achieve the compression of the PFCSR method. Since this compression method performs secondary compression on each column index on the basis of CSR, it can overall reduce the bit width of the column index, reduce the storage space occupied by the column index, and thus achieve a higher compression ratio, save storage space, and can reduce the data volume during data reading and writing, and improve the operation efficiency.
[0207] After storing the compressed second matrix in the main memory of the computing device, the cache unit in the computing device can read the compressed second matrix, which can reduce the data volume transmitted between the cache unit and the main memory. After reading the compressed second matrix, the cache unit can represent the compressed second matrix as the sparse matrix representation before compression through decompression, and then can input each element into the computing unit to participate in the operation.
[0208] An embodiment of the present application also provides an operation method.Figure 13 A schematic diagram of the operation method provided by the embodiment of the present application. The operation method is applied to the operation device in any of the above embodiments. The operation device includes a scheduling unit, a cache unit, and at least one operation unit. The method is used for the multiplication operation of the first matrix multiplied by the second matrix. As Figure 13 shown, it includes:
[0209] S1301, according to the non-zero elements in the first matrix, determine the number of multiplications of any row in the second matrix participating in the multiplication operation, and indicate the number of multiplications of each row in the second matrix to the cache unit.
[0210] S1302, according to the number of multiplications of each row, read and cache the non-zero elements of each row in the second matrix.
[0211] S1303, input each non-zero element of any row in the first matrix into an operation unit.
[0212] S1304, according to each non-zero element of any row in the first matrix, read the non-zero elements of the corresponding rows in the second matrix participating in the multiplication operation in the cache unit, and perform the multiplication combination operation of the matrix to obtain the corresponding elements in the result matrix.
[0213] Optionally, determining the number of multiplications of any row in the second matrix participating in the multiplication operation according to the non-zero elements in the first matrix includes: counting the number of any column index according to the column index of the non-zero elements in the first matrix; according to the index value of any column index, determine the number of multiplications of the row in the second matrix whose row pointer is the index value of any column index participating in the multiplication operation.
[0214] Optionally, the method further includes: for the non-zero elements of multiple rows in the first matrix, compare the sizes according to the number of column indexes of each non-zero element, and adjust the position of the non-zero element with a larger number of column indexes to the front of the row where the non-zero element is located, so as to perform the multiplication operation on the non-zero element preferentially.
[0215] Optionally, the method further includes: sorting the number of column indexes of each non-zero element in the first matrix in descending order to obtain the sorting result of each column index, and indicating the sorting result to the cache unit; reading and caching the non-zero elements of each row in the second matrix in sequence according to the sorting result.
[0216] Optionally, the method further includes: determining the number of multiplications of the row where any non-zero element in the second matrix is located as the number of multiplications of the any non-zero element; if an operation unit reads a non-zero element from the cache unit each time, subtract one from the number of multiplications of the read non-zero element; in the case where the cache unit is full and new non-zero elements need to be read and cached, by comparing the number of multiplications of the cached non-zero elements with the number of multiplications of the new non-zero elements, preferentially cache the non-zero elements with a larger number of multiplications.
[0217] Optionally, any arithmetic unit includes at least two combiners, and any combiner is used to combine partial sums of any element in the result matrix, where the partial sum is an intermediate result when the matrix gradually accumulates products for any element during the multiplication operation.
[0218] Optionally, the arithmetic unit includes two combiners, one of the two combiners is a first-level combiner, and the other combiner is a second-level combiner; the first-level combiner is used to serially combine m groups of partial sums to obtain m first-level combined results, and any group of partial sums in the m groups of partial sums includes partial sums of n rows, and the n rows respectively correspond to n rows in the second matrix, where m and n are both positive integers; the second-level combiner is used to combine the m first-level combined results to obtain a second-level combined result, and the second-level combined result is used to obtain the elements of one row in the result matrix.
[0219] Optionally, the arithmetic device includes a plurality of arithmetic units. When the second matrix includes k rows, k is a positive integer, and the product of m and n is less than k, the method further includes: calling the combiners in the plurality of arithmetic units that are in an idle state to combine partial sums of the plurality of second-level combined results to obtain at least one multi-level combined result, and the multi-level combined result is used to obtain the elements of one row in the result matrix.
[0220] Optionally, the cache unit includes p groups of buffer areas, any group of buffer areas includes q buffer slices, and any buffer slice is used to cache non-zero elements in the second matrix, and any buffer slice corresponds to an arithmetic unit, where p and q are both positive integers; the method further includes: when it is necessary to call the combiners in the plurality of arithmetic units that are in an idle state, preferentially calling the arithmetic units corresponding to the buffer slices in the same group of buffer areas.
[0221] Optionally, between the p groups of buffer areas, data paths are provided between any adjacent two groups of buffer areas, and the data paths are used to transmit the non-zero elements cached in the buffer slices.
[0222] Optionally, the arithmetic device includes a first storage unit and a second storage unit. The first storage unit is used to store the second matrix, and the second storage unit is used to store the result matrix; when the second matrix is a sparse matrix, the first storage unit is used to store the compressed second matrix obtained by compressing the second matrix, and the cache unit reads the non-zero elements of each row in the compressed second matrix from the first storage unit.
[0223] Optionally, the compressed second matrix is obtained by compressing the sparse matrix in the following manner: compressing the second matrix by compressing the sparse row CSR to obtain a preliminary compressed second matrix; for the column index of each non-zero element in any row of the preliminary compressed second matrix, dividing the column index according to the bits of a preset length to obtain at least two levels of sub-indexes, and any two adjacent levels of sub-indexes have a superior-subordinate relationship; for any level of sub-index in the at least two levels of sub-indexes, merging the sub-indexes with the same sub-index of the upper level and the same sub-index of the current level into one sub-index, and marking the correspondence between the merged sub-index and the sub-index of the same column index in the lower level sub-index; replacing the column indexes in the preliminary compressed second matrix with the merged sub-indexes and the marked correspondence to obtain a compressed second matrix.
[0224] The implementation principle and technical effect of the computing method provided in the embodiment of the present application are similar to those of the embodiments of the above-mentioned computing devices, and will not be described in detail in this embodiment.
[0225] The embodiment of the present application also provides a board card, which may be a board card in an electronic device. Figure 14 A schematic diagram of a board provided in an embodiment of the present application is shown as follows: Figure 14 As shown, the board includes a chip 1401, which can be a system-on-chip (SoC), or a system on chip, integrated with one or more computing devices, and the computing device can be any computing device of the above-mentioned embodiments of the present application. Chip 1401 can be used to support various deep learning and machine learning algorithms to meet the matrix operation and processing requirements in complex scenarios in the fields of computer vision, speech, natural language processing, data mining, etc. For example, deep learning technology is widely used in the field of cloud intelligence. One of the characteristics of cloud intelligence applications is the large amount of input data, which has high requirements on the storage capacity and computing power of the platform. The board of this embodiment is suitable for cloud intelligence applications, has a high matrix multiplication efficiency, and can meet the needs of deep learning tasks.
[0226] Optionally, the chip 1401 is connected to an external device 1403 via an external interface device 1402. The external device 1403 may be, for example, a server, a computer, a camera, a display, a mouse, a keyboard, a network card, or a wifi interface. The data to be processed may be transmitted from the external device 1403 to the chip 1401 via the external interface device 1402. The calculation result of the chip 1401 may be transmitted back to the external device 1403 via the external interface device 1402. According to different application scenarios, the external interface device 1402 may have different interface forms, such as a PCIe interface.
[0227] The board also includes a storage device 1404 for storing data, which includes one or more storage units 1405. The storage device 1404 is connected to the control device 1406 and the chip 1401 through a bus for data transmission. The control device 1406 in the board is configured to control the state of the chip 1401. For this purpose, in an application scenario, the control device 1406 may include a microcontroller unit (MCU), etc.
[0228] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0229] It should be further noted that although each step in the flowchart is shown in sequence according to the arrow indication, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0230] It should be understood that the above device embodiments are illustrative, and the devices of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0231] In addition, without special instructions, in each embodiment of this application, each functional unit / module can be integrated in one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.
[0232] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor can be any suitable hardware processor, such as CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0233] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes: USB flash drive, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk, or optical disc, etc., all kinds of media that can store program codes.
[0234] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
Claims
1. A computing device, characterized in that: The operation device is applied to a multiplication operation of a first matrix by a second matrix, and the operation device comprises a scheduling unit, a cache unit and at least one operation unit; The scheduling unit is used to determine the number of multiplications of any row in the second matrix participating in the multiplication operation according to the non-zero elements in the first matrix, and indicate the number of multiplications of each row in the second matrix to the cache unit; The cache unit is used to read and cache the non-zero elements of each row in the second matrix according to the number of multiplications of each row; The scheduling unit is further used to input each non-zero element of any row in the first matrix into one of the operation units; The operation unit is used to read the non-zero elements corresponding to each row of the second matrix participating in the multiplication operation in the cache unit according to each non-zero element in any row of the first matrix, and perform matrix multiplication and merging operations to obtain corresponding elements in the result matrix.
2. The device according to claim 1, characterized in that The scheduling unit is specifically used for: According to the column indices of the non-zero elements in the first matrix, counting the number of any column index; According to the index value of any column index, the number of any column index is determined as the number of times rows in the second matrix whose row pointers are the index value of any column index participate in the multiplication operation.
3. The device according to claim 2, characterized in that The scheduling unit is specifically used for: For the non-zero elements in multiple rows in the first matrix, a size comparison is performed based on the number of column indexes of each non-zero element, and the position of the non-zero element with a larger number of column indexes is adjusted to the front column of the row where the non-zero element is located, so as to give priority to multiplication operations on the non-zero element.
4. The device according to claim 3, characterized in that The scheduling unit is also used for: Sorting in descending order according to the number of column indexes of each non-zero element in the first matrix to obtain a sorting result of each column index, and indicating the sorting result to the cache unit; The cache unit is specifically used for: The non-zero elements of each row in the second matrix are read and cached in sequence according to the sorting result.
5. The device according to claim 1, characterized in that The cache unit is also used for: Determine the number of multiplications of the row where any non-zero element in the second matrix is located as the number of multiplications of the any non-zero element; If any of the computing units reads a non-zero element from the cache unit once, the number of multiplications of the non-zero element read is reduced by one; When the cache unit is full and a new non-zero element needs to be read and cached, the non-zero element with a larger multiplication number is cached preferentially by comparing the multiplication times of the cached non-zero elements with the multiplication times of the new non-zero elements.
6. The device according to claim 1, characterized in that Any of the operation units includes at least two mergers, and any of the mergers is used to merge the partial sum of any element in the result matrix, where the partial sum is an intermediate result when the product of any element is gradually accumulated during the multiplication operation of the matrix.
7. The device according to claim 6, characterized in that The operation unit includes two combiners, one of the two combiners is a primary combiner, and the other combiner is a secondary combiner; The first-level merger is used to merge m groups of partial sums in series to obtain m first-level merger results, wherein any group of partial sums in the m groups of partial sums includes n rows of partial sums, and the n rows respectively correspond to n rows in the second matrix, wherein m and n are both positive integers; The secondary merger is used to merge the m primary merger results to obtain a secondary merger result, and the secondary merger result is used to obtain the elements of a row in the result matrix.
8. The device according to claim 7, characterized in that The computing device includes a plurality of computing units, and when the second matrix includes k rows, k is a positive integer, and the product of m and n is less than k, the scheduling unit is further used for: The idle merger in the plurality of operation units is called to merge the partial sums of the plurality of secondary merger results to obtain at least one multi-level merger result, wherein the multi-level merger result is used to obtain an element of a row in the result matrix.
9. The device according to claim 8, characterized in that The cache unit includes p groups of cache areas, any group of the cache areas includes q cache slices, any of the cache slices is used to cache non-zero elements in the second matrix, and any of the cache slices corresponds to one of the operation units, wherein p and q are both positive integers; the scheduling unit is specifically used to: When it is necessary to call a merger in an idle state among the plurality of operation units, the operation units corresponding to the cache slices in the same group of cache areas are called preferentially.
10. The device according to claim 9, characterized in that Between the p groups of cache areas, a data path is provided between any two adjacent groups of cache areas, and the data path is used to transmit non-zero elements cached in the cache slices.
11. The device according to any one of claims 1 to 10, characterized in that: The computing device comprises a first storage unit and a second storage unit, the first storage unit is used to store the second matrix, and the second storage unit is used to store the result matrix; When the second matrix is a sparse matrix, the first storage unit is used to store a compressed second matrix obtained by sparse matrix compression of the second matrix, and the cache unit reads non-zero elements of each row in the compressed second matrix in the first storage unit.
12. The device according to claim 11, characterized in that The compressed second matrix is obtained by compressing the sparse matrix in the following manner: Compressing the second matrix by means of compressed sparse rows CSR to obtain a preliminarily compressed second matrix; For the column index of each non-zero element in any row of the preliminarily compressed second matrix, the column index is segmented according to the bits of a preset length to obtain at least two levels of sub-indexes, and any two adjacent levels of sub-indexes have a superior-subordinate relationship; For any level of sub-indexes in the at least two levels of sub-indexes, merge the sub-indexes with the same sub-index of the upper level and the same sub-index of the current level into one sub-index, and mark the corresponding relationship between the merged sub-index and the sub-indexes with the same column index in the sub-indexes of the lower level; The corresponding relationship between each merged sub-index and label is used to replace each column index in the preliminary compressed second matrix to obtain the compressed second matrix.
13. A computing method, characterized in that: The computing method is applied to the computing device according to any one of claims 1 to 12, the computing device comprising a scheduling unit, a cache unit and at least one computing unit, and the method is used for a multiplication operation of a first matrix by a second matrix, comprising: Determine the number of times any row in the second matrix participates in the multiplication operation according to the non-zero elements in the first matrix, and indicate the number of times each row in the second matrix participates in the multiplication operation to the cache unit; Read and cache non-zero elements of each row in the second matrix according to the number of multiplications of each row; Inputting each non-zero element of any row in the first matrix into one of the operation units; According to each non-zero element in any row of the first matrix, the non-zero elements corresponding to each row of the second matrix participating in the multiplication operation are read in the cache unit, and the matrix multiplication and merging operation is performed to obtain the corresponding elements in the result matrix.
14. A board, characterized in that: The board includes a computing device as described in any one of claims 1-12.
Citation Information
Patent Citations
FPGA-based graph convolutional neural network sparse matrix multiplication distribution system
CN115390788A
Sparse matrix operation method and apparatus, and computing device
CN119226681A
Method for accelerating sparse matrix calculation on tensor processing unit and storage medium
CN119441698A
Data processing method and device for sparse matrix multiplication calculation, equipment and medium
CN119848404A
Computer-implemented accumulation method for sparse matrix multiplication applications
US20240004954A1
Cited By
Data processing method and device of processor, equipment, medium and program product
CN121502136A
Data processing method, device, equipment, medium and program product of processor
CN121502136B