Computing devices, computing methods and related products
Patent Information
- Application Number
- CN202210635540.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2042-06-06
AI Technical Summary
现有的硬件不能充分利用稀疏矩阵的特性,有效地支持稀疏矩阵的相关运算
[0010] By providing the computing device configured above for performing sparse matrix multiplication, the method for using the computing device to perform sparse matrix multiplication, the chip, and the board, this disclosure provides a computing device that supports sparse matrix multiplication, which can greatly reduce computational complexity by optimizing the multiplication process, thereby improving the processing efficiency of the machine.
Smart Images

Figure CN117235424B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of data processing. More specifically, this disclosure relates to a computing device configured for performing sparse matrix multiplication, a method for performing sparse matrix multiplication using the computing device, chips, and boards. Background Technology
[0002] Sparse matrices account for a certain proportion in data processing. For example, deep learning algorithms, which have developed rapidly in recent years, are computationally intensive and storage-intensive tools. As information processing tasks become increasingly complex and the requirements for the real-time performance and accuracy of algorithms continue to increase, neural networks are often designed to be deeper and deeper, resulting in ever-increasing computational and storage requirements. This makes it difficult to directly apply existing deep learning-based artificial intelligence technologies to mobile phones, satellites, or embedded devices with limited hardware resources.
[0003] Therefore, the compression, acceleration, and optimization of deep neural network models have become particularly important. Sparsity is one such method for lightweighting models. Network parameter sparsity reduces redundant components in larger networks through appropriate methods, thereby reducing the network's computational and storage requirements. This network parameter sparsity will produce a sparse matrix.
[0004] The properties of sparse matrices allow them to significantly reduce computational and storage complexity by leveraging data structure characteristics. Therefore, research and hardware implementation of sparse matrix storage and computation methods can greatly improve the processing performance of related data structures. Existing hardware cannot fully utilize the characteristics of sparse matrices to effectively support their related operations. Summary of the Invention
[0005] In order to at least partially solve one or more of the technical problems mentioned in the background art, the present disclosure provides a computing device configured for performing sparse matrix multiplication, a method for performing sparse matrix multiplication using the computing device, a chip, and a board.
[0006] In a first aspect, this disclosure discloses a computing device configured to perform sparse matrix multiplication operations, the computing device comprising: a storage circuit storing a left multiplication matrix in row-compressed format and a right multiplication matrix in column-compressed format; and a processing circuit configured to perform a bitwise multiplication accumulation operation on element values whose column indices in the row data are the same as those in the column data, based on row data of the left multiplication matrix and column data of the right multiplication matrix from the storage circuit, to obtain element values of corresponding rows and columns in the result matrix.
[0007] In a second aspect, this disclosure provides a chip that includes a computing device according to any of the embodiments of the first aspect.
[0008] In a third aspect, this disclosure provides a board including the chip of any of the embodiments of the second aspect above.
[0009] In a fourth aspect, this disclosure provides a method for performing sparse matrix multiplication using a computing device according to any embodiment of the first aspect.
[0010] By providing the computing device configured above for performing sparse matrix multiplication, the method for using the computing device to perform sparse matrix multiplication, the chip, and the board, this disclosure provides a computing device that supports sparse matrix multiplication, which can greatly reduce computational complexity by optimizing the multiplication process, thereby improving the processing efficiency of the machine. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0012] Figure 1 This diagram shows the structure of the board card according to an embodiment of this disclosure;
[0013] Figure 2 This diagram illustrates the structure of the combined processing apparatus according to an embodiment of the present disclosure.
[0014] Figure 3 A schematic diagram illustrating the internal structure of a processor core in a single-core or multi-core computing device according to embodiments of the present disclosure;
[0015] Figure 4 Several storage methods for sparse matrices are shown;
[0016] Figure 5 A schematic structural block diagram of a computing device according to an embodiment of the present disclosure is shown; and
[0017] Figure 6 A schematic structural block diagram of a computing device according to another embodiment of this disclosure is shown. Detailed Implementation
[0018] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0019] It should be understood that the terms "first," "second," "third," and "fourth," etc., that may appear in the claims, specification, and drawings of this disclosure are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0020] It should also be understood that the terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0021] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection."
[0022] The specific embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.
[0023] Exemplary computing environment
[0024] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of this disclosure is shown. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.
[0025] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.
[0026] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).
[0027] Figure 2 This is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. (As shown) Figure 2 As shown, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203, and a storage device 204.
[0028] The computing device 201 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.
[0029] Interface device 202 is used to transmit data and control commands between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to on-chip storage device of computing device 201. Further, computing device 201 can obtain control commands from processing device 203 via interface device 202 and write them to on-chip control cache of computing device 201. Alternatively or optionally, interface device 202 can also read data from storage device of computing device 201 and transmit it to processing device 203.
[0030] Processing device 203, as a general-purpose processing device, performs basic control including but not limited to data transfer, and starting and / or stopping computing device 201. Depending on the implementation, processing device 203 may be one or more types of processors, including but not limited to digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, computing device 201 disclosed herein can be considered as having a single-core structure or a homogeneous multi-core structure. However, when computing device 201 and processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.
[0031] Storage device 204 is used to store data to be processed. It may be DRAM or DDR memory, typically 16G or larger in size, and is used to store data of computing device 201 and / or processing device 203.
[0032] Figure 3 The diagram shows the internal structure of the processor core when the computing device 201 is a single-core or multi-core device. The computing device 301 is used to process input data such as computer vision, speech, natural language, and data mining. The computing device 301 includes three main modules: a control module 31, an arithmetic module 32, and a storage module 33.
[0033] The control module 31 coordinates and controls the operation of the computation module 32 and the storage module 33 to complete the deep learning task. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 fetches instructions from the processing device 203, and the instruction decode unit 312 decodes the fetched instructions and sends the decoding result as control information to the computation module 32 and the storage module 33.
[0034] The computation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit 322 is responsible for the core computations of deep learning algorithms, namely matrix multiplication and convolution.
[0035] Storage module 33 is used to store or move relevant data, including neuron RAM (NRAM) 331, weight RAM (WRAM) 332, and direct memory access (DMA) module 333. NRAM 331 is used to store input neurons, output neurons, and intermediate results after computation; WRAM 332 is used to store the convolution kernels of the deep learning network, i.e., the weights; DMA 333 is connected to DRAM 204 through bus 34 and is responsible for data transfer between computing device 301 and DRAM 204.
[0036] Exemplary sparse matrix storage method
[0037] In a matrix, if the number of elements with the value 0 far exceeds the number of non-zero elements, and the distribution of non-zero elements is irregular, then the matrix is called a sparse matrix. Simply put, a sparse matrix is a matrix where the vast majority of elements are 0, containing only a small number of non-zero values. Below is an example of a sparse matrix of size 3×6:
[0038]
[0039] Since sparse matrices are mostly composed of zero elements, storing them in the usual way would undoubtedly waste a lot of space. At the same time, if calculations are performed in the usual way, the zero elements do not help the final result and instead increase a lot of unnecessary calculations.
[0040] To achieve compression, sparse matrix storage only stores the non-zero element values (sometimes called valid element values), but also retains the positions of the non-zero elements for easy recovery. Therefore, sparse matrix storage not only stores the non-zero element values but also their coordinate positions (row index and column index).
[0041] Figure 4 Several storage methods for sparse matrices are shown. Or, in other words, Figure 4 Several data structures are provided for storing sparse matrices.
[0042] The COO storage method, also known as the coordinate format, uses three arrays to store the sparse matrix. These arrays store the row index (row number), column index (column number), and value of the non-zero elements. The length of each array is the number of non-zero elements in the sparse matrix. Theoretically, the elements in the sparse matrix can be stored in any order, but for ease of access, they are stored in left-to-right, top-to-bottom order—that is, in row-major order.
[0043] As shown in the figure, for the first element "1", its row index is 0, its column index is 0, and its value is 1; for the second element "2", its row index is 0, its column index is 1, and its value is 2; other non-zero elements are stored similarly. It can be seen from the figure that in this COO storage method, there are some duplicate elements in the row index array and the column index array.
[0044] CSR storage, also known as Compressed Sparse Row Format, compresses the row index array in the COO format while leaving the other two arrays unchanged. This results in three arrays: row pointers, column indices, and values. The lengths of the column index array and the value array remain the same, counting the number of non-zero elements. The row pointer array stores the offset of the first non-zero element in each row from the first non-zero element in the sparse matrix, and its last element stores the total number of non-zero elements in the sparse matrix. Therefore, the length of the row pointer array is the number of rows in the sparse matrix plus one.
[0045] As shown in the diagram, the column index array and the numeric array are the same as in the COO method. For the row pointer array, the first element is "0", which means that the first non-zero element "1" in row 0 is offset by 0 from the first non-zero element in the sparse matrix, since it is the first non-zero element itself. The second element "2" means that the first non-zero element "2" in row 1 is offset by 2 from the first non-zero element "1" in the sparse matrix, which is equal to the number of non-zero elements in row 0. The third element "4" means that the first non-zero element "5" in row 2 is offset by 4 from the first non-zero element "1" in the sparse matrix, which is equal to the number of non-zero elements in rows 0 and 1. The fourth element "6" represents the total number of all non-zero elements.
[0046] CSC storage, also known as Compressed Sparse Column Format, is similar to CSR, except that columns are compressed while row indices and the data array remain unchanged. The column pointer array stores the offset of the first non-zero element of each column from the first non-zero element of the sparse matrix, and its last element stores the total number of non-zero elements in the sparse matrix. Therefore, the length of the column pointer array is the number of columns in the sparse matrix plus 1.
[0047] As shown in the figure, the row index array and the value array are the same as in the COO method. For the column pointer array, the first element is "0", which represents the first non-zero element in column 0. The second element "1" indicates that the first non-zero element "2" in column 1 is offset by 1 from the first non-zero element "1" in the sparse matrix, which is equal to the number of non-zero elements in column 0; the third element "3" indicates that the first non-zero element "4" in column 2 is offset by 3 from the first non-zero element "1" in the sparse matrix, which is equal to the number of non-zero elements in columns 0 and 1, and so on, to obtain the other column pointer elements.
[0048] In the embodiments disclosed herein, when a sparse matrix is provided in any of the above or other compressed storage methods, it can be referred to as a compact form of the sparse matrix, or a matrix form consisting of valid data elements (e.g., non-zero values) in the sparse matrix can be referred to as a compact matrix. When a sparse matrix is provided in a compact form, the compact matrix can be partitioned in the row or column direction for subsequent processing.
[0049] For example, in one embodiment of this disclosure, the matrix can be partitioned in the row direction according to the row-wise compressed format (e.g., CSR) storage method of the left-multiplied matrix, thereby outputting the partitioned row vectors and corresponding column indices for processing. Still using... Figure 4 Taking a matrix as an example, after row partitioning of a compact matrix in CSR format consisting of valid data elements (1,2,3,4,5,6), three row vectors (1,2), (3,4) and (5,6) are output. Each row vector has a corresponding column index vector (0,1), (1,2) and (3,5), which respectively indicate the column index of each data element in the corresponding row vector.
[0050] For example, in one embodiment of this disclosure, the matrix can be partitioned in the column direction according to the column-compressed format (e.g., CSC) of the right-multiplied matrix, thereby outputting the partitioned column vectors and corresponding row indices for processing. Still using... Figure 4 Taking the matrix as an example, after the compact matrix in CSR format composed of valid data elements (1,2,3,4,5,6) is column-splittered, five column vectors (1), (2,3), (4), (5) and (6) are output. Each column vector has a corresponding row index vector (0), (0,1), (1), (2) and (2), which respectively indicate the row index of each data element in the corresponding column vector.
[0051] Exemplary sparse matrix operation principle
[0052] The principle of sparse matrix multiplication operation used in the embodiments disclosed herein is described below.
[0053] According to the definition of matrix multiplication, matrix multiplication can be calculated using formula (1):
[0054]
[0055] As can be seen from formula (1), the element in the i-th row and j-th column of the resulting matrix is obtained by multiplying and summing the elements with the same column index in the i-th row of the left matrix and the same row index in the j-th column of the right matrix. That is, the element cij in the resulting matrix can be represented as:
[0056] cij=∑ k aik·bkj (2)
[0057] The above operation method also applies to sparse matrix multiplication. Formula (3) shows the calculation formula when both the left and right multiplication matrices are sparse matrices.
[0058]
[0059]
[0060] As can be seen from the above formula, when any element in the i-th row of the left-multiplied matrix has the same column index as the j-th column of the right-multiplied matrix and is zero, the product does not contribute to the final result. In other words, multiplication and accumulation are only required when the column index of the non-zero element in the i-th row of the left-multiplied matrix is the same as the row index of the non-zero element in the j-th column of the right-multiplied matrix.
[0061] Therefore, in this disclosed embodiment, by adding a judgment on whether the column index of the non-zero element in the i-th row of the left multiplication matrix is the same as the row index of the non-zero element in the j-th column of the right multiplication matrix, the multiplication-accumulation operation can be selectively performed, thereby saving computing resources and avoiding invalid calculations.
[0062] Exemplary computing device
[0063] Figure 5 A schematic structural block diagram of a computing device 500 according to an embodiment of this disclosure is shown. It can be understood that this structure can be considered as based on... Figure 3 One implementation method is a single processing core, which can also be viewed as multiple... Figure 3 The diagram illustrates a combined implementation based on the processing core.
[0064] As shown in the figure, the computing device 500 can be configured to perform matrix multiplication operations, including but not limited to sparse matrix multiplication operations, and may include a storage circuit 510 and a processing circuit 530.
[0065] Storage circuit 510 stores left-multiplied matrices in row-compressed (e.g., CSR) format and right-multiplied matrices in column-compressed (e.g., CSC) format. As mentioned earlier, the CSR format includes three parameters: row pointer, column index, and value. Therefore, the row vectors of the left-multiplied matrix and the corresponding column index vectors can be extracted according to the CSR format. The CSC format includes three parameters: column pointer, row index, and value. Therefore, the column vectors of the right-multiplied matrix and the corresponding row index vectors can be extracted according to the CSC format.
[0066] The processing circuit 530 can be used to perform a positional multiplication and accumulation operation on elements whose column indices in the row data are the same as those in the column data, based on the row data of the left-multiplied matrix and the column data of the right-multiplied matrix from the storage circuit, to obtain the element values of the corresponding row and column in the result matrix.
[0067] In one embodiment, the processing circuit 530 may include a comparison circuit 531 and a multiply-accumulate circuit 532. The comparison circuit 531 is used to determine whether there are identical index elements in the column indices of the row data and the row indices of the column data, and outputs a corresponding indication. The multiply-accumulate circuit 532 is used to perform multiplication operations on the row elements and column elements with the same index according to the indication of the comparison circuit, and accumulate the products belonging to the same result matrix elements.
[0068] In the above scheme, by using different compression formats for left-multiplying and right-multiplying matrices, the column and row index information provided in the compression format can be used to determine whether multiplication and accumulation operations need to be performed, thereby avoiding invalid calculations and improving computational efficiency.
[0069] With the development of hardware technology, modern intelligent processors mostly adopt multi-core / multi-processor parallel architectures. Therefore, in some embodiments disclosed herein, a scheme is further provided to improve the efficiency of matrix multiplication operations when multiple parallel processing circuits are used.
[0070] Figure 6 A schematic structural block diagram of a computing device 600 according to another embodiment of this disclosure is shown. In this example, the computing device utilizes a master-slave processing circuit to perform the matrix multiplication operation described above. As shown, the computing device 600 includes a first storage circuit 610, a second storage circuit 620, a master processing circuit (MA) 630, and a plurality of slave processing circuits (SL) 640, of which 16 slave processing circuits SL0 to SL15 are schematically shown in the figure. Those skilled in the art will understand that the number of slave processing circuits can be more or less, depending on the specific hardware configuration, and the embodiments of this disclosure are not limited in this respect.
[0071] The master processing circuit and slave processing circuits, as well as multiple slave processing circuits, can communicate with each other through various connections. In different application scenarios, the connection between multiple slave processing circuits can be either a hard connection arranged by hardwired lines or a logical connection configured according to, for example, microinstructions, to form a topology of various slave processing circuit arrays. The embodiments disclosed herein are not limited in this respect. The master processing circuit and slave processing circuits can cooperate with each other to achieve parallel processing.
[0072] To support computational functions, the main processing circuit and the slave processing circuit can include various computing circuits, such as vector operation units and matrix operation units. The vector operation unit is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit is responsible for the core computations of deep learning algorithms, such as matrix multiplication and convolution.
[0073] The processing circuit can, for example, perform intermediate operations on the corresponding data in parallel according to the operation instructions to obtain multiple intermediate results, and then transmit the multiple intermediate results back to the main processing circuit.
[0074] By setting the processing circuit in a master-slave structure (e.g., a master-multiple-slave structure, or a multi-master-multiple-slave structure, which is not limited in this disclosure), for forward calculation instructions, the data can be split according to the calculation instructions, so that the computationally intensive parts can be performed in parallel by multiple slave processing circuits to improve the calculation speed, save calculation time, and thus reduce power consumption.
[0075] The first storage circuit 610 can be used to store multicast data, meaning that the data in the first storage circuit will be transmitted to multiple slave processing circuits via a broadcast bus, and these slave processing circuits will receive the same data. It can be understood that broadcasting and multicasting can be implemented via a broadcast bus. Multicast refers to a communication method that transmits a single data set to multiple slave processing circuits; while broadcasting is a communication method that transmits a single data set to all slave processing circuits, and is a special case of multicast. Since both multicast and broadcasting correspond to one-to-many transmission methods, this document does not specifically distinguish between the two; broadcasting and multicast can be collectively referred to as multicast, and those skilled in the art can understand their meaning from the context.
[0076] The second storage circuit 620 can be used to store and distribute data, that is, the data in the second storage circuit will be transmitted to different slave processing circuits respectively, and each slave processing circuit receives different data.
[0077] It is understood that the first storage circuit and the second storage circuit can be two storage blocks formed by partitioning the same memory, or they can be two independent memories; no specific limitation is made here. By providing the first storage circuit and the second storage circuit respectively, it is possible to support the transmission of data to be processed in different transmission methods, thereby reducing the amount of data access by multiplexing multicast data among multiple slave processing circuits. Either the row data of the left multiplication matrix or the column data of the right multiplication matrix can be determined as broadcast data and stored in the first storage circuit for broadcast transmission to each slave processing circuit before and / or during the operation; at this time, the other of the row data of the left multiplication matrix and the column data of the right multiplication matrix can be determined as distribution data and stored in the second storage circuit for distribution to the corresponding slave processing circuit before and / or during the operation.
[0078] In some embodiments, the left multiplication matrix can be determined as multicast data and stored in a first storage circuit to transmit the data via broadcast to multiple scheduled slave processing circuits during computation. The left multiplication matrix uses a row-compressed format, such as CSR format. Therefore, the row data can be broadcast row by row (e.g., split into row vectors). Correspondingly, the right multiplication matrix can be determined as distribution data and stored in a second storage circuit. This distribution data can be distributed to the corresponding slave processing circuits before computation. The right multiplication matrix uses a column-compressed format, such as CSC format. Therefore, the column data can be distributed column by column.
[0079] In other embodiments, the right-multiplication matrix can be defined as multicast data and stored in a first storage circuit to transmit the data via broadcast to multiple scheduled slave processing circuits during computation. The right-multiplication matrix uses a column-compressed format. Therefore, the column data can be broadcast column-wise (e.g., split into column vectors). Correspondingly, the left-multiplication matrix can be defined as distribution data and stored in a second storage circuit. This distribution data can be distributed to the corresponding slave processing circuits before computation. The left-multiplication matrix uses a row-compressed format. Therefore, the row data can be distributed row-wise.
[0080] As can be seen from the preceding principle description, matrix multiplication has a certain degree of data locality. For example, the value of the element in the i-th row and j-th column of the result matrix is determined by the i-th row of the left-multiplied matrix and the j-th column of the right-multiplied matrix. Therefore, this characteristic can be reasonably utilized to achieve parallel computing.
[0081] In some embodiments, the result matrix can be split according to its columns, distributing the operations of each column to different slave processing circuits. That is, different columns of right-multiplication matrices can be assigned to different slave processing circuits, while the same left-multiplication matrix row data is used for the operation, corresponding to the row data broadcasting and column data distribution operation mode described above. In this case, the left-multiplication matrix row data is reused among these processing circuits, with the number of reuses equal to the number of processing circuits. It can be understood that when the number of columns in the result matrix exceeds the number of schedulable slave processing circuits, it needs to be split into multiple rounds of operation to complete the overall operation. In this operation mode, multiple slave processing circuits can cyclically and in parallel calculate the column data of adjacent different columns in the result matrix. For example, assuming there are N slave processing circuits, each of which can process one column of result data at a time, then in one calculation loop, the N slave processing circuits can concurrently calculate the first N columns of data; after calculation, in the next loop, the N slave processing circuits can concurrently calculate the next N columns of data.
[0082] In other embodiments, the result matrix can be split according to its rows, distributing the operations of each row to different slave processing circuits. That is, different rows of left-multiplication matrices can be assigned to different slave processing circuits, while the same right-multiplication matrix column data is used for the operation, corresponding to the column data broadcasting and row data distribution operation mode described above. In this case, the right-multiplication matrix column data is reused among these processing circuits, with the number of reuses equal to the number of processing circuits. It can be understood that when the number of rows in the result matrix exceeds the number of schedulable slave processing circuits, it needs to be split into multiple rounds of operation to complete the overall operation. In this operation mode, multiple slave processing circuits can cyclically and in parallel calculate the row data of adjacent different rows in the result matrix. For example, assuming there are N slave processing circuits, and each slave processing circuit can process one row of result data at a time, then in one calculation loop, the N slave processing circuits can concurrently calculate the first N rows of data; after calculation, in the next loop, the N slave processing circuits can concurrently calculate the next N rows of data.
[0083] In some cases, due to limitations in processing bandwidth or hardware capabilities, the amount of data that a processing circuit can process at one time is limited. If the row vectors of the left-multiplied matrix or the column vectors of the right-multiplied matrix are very large, they cannot be processed in one go and need to be divided into multiple processing steps.
[0084] For example, broadcast data is divided into multiple broadcast vectors according to the index order within the same row (for left-multiplied matrix row data broadcast) or the same column (for right-multiplied matrix column data broadcast), and broadcast sequentially to the multiple scheduled slave processing circuits. Correspondingly, distribution data is divided into multiple distribution vectors according to the index order within the same column (for right-multiplied matrix column data distribution) or the same row (for left-multiplied matrix row data distribution), and distributed sequentially to the corresponding slave processing circuits.
[0085] Within the processing circuit, the indices of the broadcast vector and the distribution vector need to be compared. Elements with the same index are then multiplied and accumulated bitwise. After processing the current broadcast vector, the processing circuit can request to broadcast the next broadcast vector. Considering that the indices of the broadcast vector and the distribution vector are ordered, for example, in ascending order, in some embodiments, the next broadcast vector can be broadcast if the largest index in the current broadcast vector is less than the smallest index among all the currently distributed data (i.e., the distribution vectors in each processing circuit). In this case, it means that it is impossible to find an index in the current broadcast vector that is the same as any current distribution vector; that is, there is no overlapping area between the indices, therefore, it is necessary to search for the next broadcast vector.
[0086] On the other hand, once the current distribution vector has been processed, the processing circuit can request the distribution of the next distribution vector. In some embodiments, the next distribution vector can be distributed when the largest index in the current distribution vector in the processing circuit is less than the smallest index in the current broadcast vector. In this case, it means that it is impossible to find an index in the current distribution vector that is the same as the current broadcast vector, and the next distribution vector needs to be searched.
[0087] Therefore, by updating the broadcast vector and distribution vector under certain conditions, the specified row and column vectors can be processed. Then, the next row and / or the next column of data can be changed, which will not be repeated here.
[0088] The computing device may also include an output storage circuit (not shown) for storing the result matrix output from the processing circuit. The output storage circuit may reuse the first storage circuit or the second storage circuit, or it may be a separate storage circuit; the embodiments disclosed herein are not limited in this respect.
[0089] Figure 6 A schematic diagram of the internal structure of the slave processing circuit SL according to an embodiment of this disclosure is also shown. As shown, each slave processing circuit 640 may include one or more arithmetic circuits CU 641. Four arithmetic circuits CU0 to CU3 are shown in the figure. Those skilled in the art will understand that the number of arithmetic circuits may be more or less, depending on the specific hardware configuration, and the embodiments of this disclosure are not limited in this respect.
[0090] Each arithmetic circuit CU 641 may include a comparator 651 and a multiply-accumulate unit 652. The comparator 651 compares the column index of the left multiplication row data assigned to the corresponding arithmetic circuit with the row index of the right multiplication column data, identifies the elements with the same index, and outputs a corresponding indication. The multiply-accumulate unit 652 performs multiplication operations on row and column elements with the same index according to the indication from the comparator 651, and accumulates the products belonging to the same result matrix elements.
[0091] In some embodiments, the slave processing circuit may further include a first buffer circuit 642 and a second buffer circuit 643. The first buffer circuit 642 may be used to buffer one or more row data of a left-multiplied matrix allocated to the slave processing circuit. Since the left-multiplied matrix uses a row-compressed format, such as CSR format, the row data includes a row pointer, a column index of the corresponding valid element in the row, and the element value. Correspondingly, the second buffer circuit 643 may be used to buffer one or more column data of a right-multiplied matrix allocated to the slave processing circuit. Since the right-multiplied matrix uses a column-compressed format, such as CSC format, the column data includes a column pointer, a row index of the corresponding valid element in the column, and the element value.
[0092] In some embodiments, the result matrix can be split according to its columns, and the operations on each column can be distributed to different computational circuits. For example, when there are 16 processing circuits, each including 4 computational circuits, the total number of schedulable computational circuits is 64. The 64 columns of data in the right-multiplication matrix can be sequentially distributed to these 64 computational circuits for computation. After processing is completed, the next 64 columns of data can be read and processed. Thus, in some embodiments, the row data in the first buffer circuit can be broadcast to multiple computational circuits; the column data in the second buffer circuit can be distributed to the corresponding computational circuits column by column.
[0093] In other embodiments, the resulting matrix can be split according to its rows, distributing the operations of each row to different computational circuits. The specific splitting is similar to column splitting and will not be repeated here.
[0094] The processing circuit 640 may also include a third buffer circuit 644 for buffering the calculation results of each arithmetic circuit CU 641.
[0095] Understandable, although Figure 6 The various processing circuits and storage circuits are shown as separate modules, but depending on the configuration, the storage circuits and processing circuits can also be combined into a single module. For example, the first storage circuit 610 can be combined with the main processing circuit 630, while the second storage circuit 620 can be shared by multiple slave processing circuits 640, with each slave processing circuit allocated an independent storage area to accelerate access. This disclosure does not limit the embodiments in this respect. Furthermore, in this computing device, the main processing circuit and slave processing circuits can belong to different modules of the same processor or chip, or they can belong to different processors; this disclosure also does not limit this in this respect.
[0096] The comparator 651 in the operational circuit can be implemented in various ways. The following describes a processing scheme for a broadcast vector and a distribution vector.
[0097] In some embodiments, comparator 651 includes a vector comparator that polls the index elements of the distribution vector based on the broadcast vector to find identical index elements. For example, the vector comparator can poll the row index elements in the row index vector of the column data distributed to the corresponding arithmetic circuit based on the column index vector of the broadcast row data to find the row index element that is the same as any column index element in the column index vector. Specifically, the first row index element can be copied and extended to the corresponding length and compared with the column index vector of the row data. If the output is all zeros, it means that there is no identical element, so the second row index element is processed. If there is a 1 in the output, it means that there is an identical element, and the multiply-accumulate unit can be instructed to perform a multiply-accumulate operation on the corresponding element value; and so on, until the row index vector of the column data has been traversed. When broadcasting column data, the scheme is similar and will not be repeated here.
[0098] In other embodiments, comparator 651 includes a vector comparator that compares the row index vector of the column data distributed to the corresponding arithmetic circuit with the row index elements in the column index vector of the broadcast row data to find the same index elements. It can be seen that in this embodiment, only the row index element of one point is broadcast at a time.
[0099] In some other embodiments, comparator 651 may include a shift register in addition to a vector comparator. One of the column index vector and the row index vector of the row data can be used as the comparison base in the vector comparator, and the other vector is stored in the shift register. The register is shifted one bit at a time and compared with the comparison base to find the common index elements in the column index vector and the row index vector.
[0100] Furthermore, considering that the index elements in an index vector are usually arranged in ascending order, when comparing by shifting, the index vector with the largest first index element can be set as the comparison benchmark, thereby avoiding possible omissions. Those skilled in the art will also understand that, based on this ordered arrangement of index elements, the comparison process can be optimized to speed up processing.
[0101] As described above, this disclosure provides a computing device that effectively supports sparse matrix multiplication. Furthermore, this disclosure provides a comparison circuit that performs operations on the non-zero element values in the sparse matrix; all operations are valid, thus eliminating wasted computing power. Even further, this disclosure can be implemented on multiple parallel processing circuits, thereby improving machine processing efficiency. In addition, some embodiments of this disclosure also provide methods, chips, and boards for performing sparse matrix multiplication using this computing device, which include features corresponding to those described above for the computing device, and will not be repeated here.
[0102] Depending on the application scenario, the electronic devices or apparatus disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The electronic devices or apparatus disclosed herein can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the electronic devices or apparatus disclosed herein can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as cloud computing, edge computing, and terminal applications. In one or more embodiments, the high-computing-power electronic devices or apparatuses according to the present disclosure can be applied to cloud devices (e.g., cloud servers), while the low-power electronic devices or apparatuses can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration.
[0103] It should be noted that, for the sake of brevity, this disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art will understand that the solutions disclosed herein are not limited by the order of the described actions. Therefore, based on the disclosure or teachings of this document, those skilled in the art will understand that some steps can be performed in a different order or simultaneously. Furthermore, those skilled in the art will understand that the embodiments described in this disclosure can be considered optional embodiments, that is, the actions or modules involved are not necessarily essential for the implementation of one or more solutions disclosed herein. In addition, depending on the solution, the description of some embodiments in this disclosure may have different emphases. In view of this, those skilled in the art will understand that parts not described in detail in a certain embodiment of this disclosure can also be referred to the relevant descriptions of other embodiments.
[0104] In terms of specific implementation, based on the disclosure and teachings of this document, those skilled in the art will understand that several embodiments disclosed herein can also be implemented in other ways not disclosed herein. For example, regarding the various units in the electronic device or apparatus embodiments described above, this document has divided them based on logical functions, but in actual implementation, there may be other ways of division. As another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed above in conjunction with the accompanying drawings can be direct or indirect couplings between units or components. In some scenarios, the aforementioned direct or indirect couplings involve communication connections utilizing interfaces, where the communication interface can support electrical, optical, acoustic, magnetic, or other forms of signal transmission.
[0105] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network units. Furthermore, depending on actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this disclosure. Additionally, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically independently.
[0106] In other implementation scenarios, the integrated units described above can also be implemented in hardware, i.e., as specific hardware circuits, which may include digital circuits and / or analog circuits. The physical implementation of the circuit's hardware structure may include, but is not limited to, physical devices, which may include, but are not limited to, transistors or memristors. Therefore, the various devices described herein (e.g., computing devices or other processing devices) can be implemented using appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage units or storage devices can be any suitable storage medium (including magnetic storage media or magneto-optical storage media), such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM.
[0107] The foregoing can be better understood in accordance with the following terms:
[0108] Clause 1. A computing device configured to perform sparse matrix multiplication operations, the computing device comprising:
[0109] A storage circuit that stores left-multiplied matrices in row-compressed format and right-multiplied matrices in column-compressed format; and
[0110] The processing circuit is configured to perform a bitwise multiplication and accumulation operation on element values whose column indices in the row data are the same as those in the row data, based on the row data of the left-multiplied matrix and the column data of the right-multiplied matrix from the storage circuit, to obtain the element values of the corresponding rows and columns in the result matrix.
[0111] Clause 2. The computing device according to Clause 1, wherein the processing circuitry includes a plurality of slave processing circuitry, the storage circuitry includes a first storage circuitry and a second storage circuitry, and one of the row data and the column data is determined as broadcast data, the broadcast data being stored in the first storage circuitry and broadcast to each slave processing circuitry, and the other of the row data and the column data is determined as distribution data, the distribution data being stored in the second storage circuitry and distributed to the corresponding slave processing circuitry.
[0112] Clause 3. The computing device according to Clause 2, wherein the row data is broadcast row by row, the column data is distributed column by column, and the plurality of slave processing circuits are used to cyclically and in parallel compute column data of adjacent different columns in the result matrix.
[0113] Clause 4. The computing device according to Clause 2, wherein the column data is broadcast column by column, the row data is distributed row by row, and the plurality of slave processing circuits are used to cyclically and in parallel compute row data of adjacent different rows in the result matrix.
[0114] Clause 5. The computing device according to any one of Clauses 2-4, wherein the broadcast data is divided into one or more broadcast vectors in the order of indexes in the same row or column and broadcast sequentially to the plurality of slave processing circuits, and the next broadcast vector is broadcast when the largest index in the current broadcast vector is less than the smallest index in the currently distributed data in all slave processing circuits.
[0115] Clause 6. The computing device according to any one of Clauses 2-5, wherein the distributed data is divided into one or more distribution vectors according to the index order in the same column or row, and distributed sequentially to the corresponding slave processing circuits, wherein the next distribution vector is distributed when the maximum index in the current distribution vector in the slave processing circuit is less than the minimum index in the current broadcast vector.
[0116] Clause 7. A computing device according to any one of Clauses 2-6, wherein each slave processing circuit includes:
[0117] A comparator is used to compare the column index of the row data assigned to the slave processing circuit with the row index of the column data, find the same index elements, and output the corresponding indication; and
[0118] A multiply-accumulate operator is used to perform multiplication operations on row and column elements with the same index, according to the instructions of the comparator, and to accumulate the products belonging to the same element in the result matrix.
[0119] Clause 8. The computing apparatus according to Clause 7, wherein the comparator includes a vector comparator configured to poll the row index elements of the column data based on the column index vector of the row data to find common index elements therein; or to poll the column index elements of the row data based on the row index vector of the column data to find common index elements therein.
[0120] Clause 9. The computing apparatus according to Clause 7, wherein the comparator includes a vector comparator and a shift register, the vector comparator being configured to perform a shift comparison with another vector stored in the shift register, using one of the column index vectors and the row index vectors of the column data as a reference, to find the same index elements therein.
[0121] Clause 10. The computing apparatus according to Clause 8, wherein the index vector whose first index element is larger in the column index vector and the row index vector is set as the reference.
[0122] Clause 11. A chip comprising a computing device according to any one of Clauses 1-10.
[0123] Clause 12, a board including the chip described in Clause 11.
[0124] Clause 13. A method for performing sparse matrix multiplication using any of the computing devices described in Clauses 1-10.
[0125] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. A computing device configured to perform sparse matrix multiplication operations, the computing device comprising: The storage circuit stores left-multiplied matrices in row-compressed format and right-multiplied matrices in column-compressed format. as well as The processing circuit is configured to perform a bitwise multiplication and accumulation operation on element values whose column indices in the row data are the same as those in the row data, based on the row data of the left-multiplied matrix and the column data of the right-multiplied matrix from the storage circuit, to obtain the element values of the corresponding row and column in the result matrix. The processing circuit includes multiple slave processing circuits, the storage circuit includes a first storage circuit and a second storage circuit, and one of the row data and the column data is determined as broadcast data, the broadcast data is stored in the first storage circuit and broadcast to each slave processing circuit, and the other of the row data and the column data is determined as distribution data, the distribution data is stored in the second storage circuit and distributed to the corresponding slave processing circuit; Each of the processing circuits includes: A comparator is used to compare the column index of the row data assigned to the slave processing circuit with the row index of the column data, find the same index element, and output the corresponding indication; as well as A multiply-accumulate operator is used to perform multiplication operations on row and column elements with the same index, according to the instructions of the comparator, and to accumulate the products belonging to the same element in the result matrix.
2. The computing device of claim 1, wherein the row data is broadcast row by row, the column data is distributed column by column, and the plurality of slave processing circuits are used to cyclically and in parallel compute column data of adjacent different columns in the result matrix.
3. The computing device of claim 1, wherein the column data is broadcast column by column, the row data is distributed row by row, and the plurality of slave processing circuits are used to cyclically and in parallel compute row data of adjacent different rows in the result matrix.
4. The computing device according to any one of claims 1-3, wherein the broadcast data is divided into one or more broadcast vectors according to the index order in the same row or column, and broadcast sequentially to the plurality of slave processing circuits, and when the largest index in the current broadcast vector is less than the smallest index in the currently distributed data in all slave processing circuits, the next broadcast vector is broadcast.
5. The computing device according to any one of claims 1-3, wherein the distributed data is divided into one or more distribution vectors according to the index order in the same column or row, and distributed sequentially to the corresponding slave processing circuits, and when the maximum index in the current distribution vector in the slave processing circuit is less than the minimum index in the current broadcast vector, the next distribution vector is distributed.
6. The computing apparatus according to any one of claims 1-3, wherein the comparator includes a vector comparator configured to poll the row index elements of the column data based on the column index vector of the row data to find identical index elements therein; or to poll the column index elements of the row data based on the row index vector of the column data to find identical index elements therein.
7. The computing device according to any one of claims 1-3, wherein the comparator comprises a vector comparator and a shift register, the vector comparator being configured to perform a shift comparison with another vector stored in the shift register, using one of the column index vector and the row index vector of the column data as a reference, to find the same index element therein.
8. The computing apparatus of claim 6, wherein the index vector whose first index element is larger in the column index vector and the row index vector is set as the reference.
9. A chip comprising a computing device according to any one of claims 1-8.
10. A circuit board comprising the chip according to claim 9.
11. A method for performing sparse matrix multiplication using the computing device according to any one of claims 1-8.
Citation Information
Patent Citations
Matrix operation-based parallel computing method
CN101980182A
Calculation engine and electronic equipment
CN106126481A
Integrated circuit chip device and related product
CN110197270A