Matrix calculation circuit, method, electronic device and computer-readable storage medium
By designing a matrix calculation circuit, using a bitmap matrix to represent the non-0 data location of the sparse matrix, the problems of low computing efficiency and high circuit complexity in the prior art are solved, and efficient sparse matrix calculation is realized.
Patent Information
- Application Number
- CN202011030120.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-09-27
AI Technical Summary
When processing sparse matrices, the prior art cannot effectively utilize sparse characteristics to accelerate matrix calculations or has problems with high circuit complexity.
By designing a matrix calculation circuit, including an instruction decoding circuit, a data reading circuit, a position data conversion circuit and a calculation circuit, the bitmap matrix represents the non-0 data position of the sparse matrix, and generates a read control signal for data reading and calculation, achieving efficient operation of the sparse matrix.
It improves the efficiency of matrix operations, reduces circuit complexity and storage costs, and is suitable for handling large-scale sparse matrix computing requirements.
Smart Images

Figure CN114282158B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of processors, and in particular, to a matrix calculation circuit, method, electronic device, and computer-readable storage medium. Background Art
[0002] With the development of science and technology, human society is rapidly entering the intelligent era. An important feature of the intelligent era is that people obtain an increasing variety and quantity of data, while having an increasingly high requirement for the speed of processing data. A chip is the cornerstone of task allocation, which fundamentally determines people's data processing ability. From the perspective of application fields, there are mainly two routes for chips: one is the general-purpose chip route, such as CPU (Central Processing Unit), etc., which can provide great flexibility, but have relatively low effective computing power when processing algorithms in specific fields; the other is the dedicated chip route, such as TPU (Tensor Processing Unit), etc., which can exert relatively high effective computing power in certain specific fields, but have poor processing ability or even cannot process in the face of flexible and diverse general fields. Due to the large variety and huge quantity of data in the intelligent era, it is required that the chip has both extremely high flexibility to handle algorithms in different fields and ever-changing algorithms, and extremely strong processing ability to quickly process a huge and rapidly growing amount of data.
[0003] In neural network computing, convolutional computing accounts for most of the total computing volume, and convolutional computing can be converted into matrix multiplication computing. Therefore, to improve the throughput, reduce latency, and enhance the effective computing power of the chip in neural network tasks, the key lies in improving the speed of matrix multiplication computing.
[0004] Many matrices composed of data in neural networks (here the data includes parameter data and input data in neural networks) are sparse matrices, that is, there are a large number of elements in the matrix with a value of 0. In order to reduce the storage amount and bandwidth occupancy of data in neural network computing, the sparse matrix will be compressed for storage; in order to improve the matrix operation speed, the sparse matrix operation will be optimized.
[0005] Figure 1a It is a schematic diagram of matrix multiplication computing in a neural network. As Figure 1a shown, M1 is a data matrix, M2 is a parameter matrix, and M is an output matrix. One row of data in M1 and one column of parameters in M2 are multiplied and added to obtain one data in M. Among them Figure 1a for the two matrices M1 and M2, one of them may be a sparse matrix, or both may be sparse matrices.
[0006] As Figure 1bThe following is a schematic diagram of matrix compression. For the storage of sparse matrices, a general compression method can be adopted: only store non-zero elements. While storing the values of these non-zero elements, the position information of the elements in the matrix will also be stored, that is, the relative coordinates X and Y of the elements in the matrix. Where X represents the row number of the matrix, and Y represents the column number of the matrix. This method takes data and coordinates as a data structure and stores them in units of this data structure. As Figure 1b shown, taking an MxN matrix as an example, it is compressed from the left MxN matrix to the right compressed matrix. Each data structure in the compressed matrix represents the non-zero data in the left matrix and the coordinates of this non-zero data in the matrix.
[0007] In a sparse matrix, since some elements in the matrix have a value of 0 and these 0 elements do not need to be stored, adopting this compression method can effectively reduce the storage capacity of the matrix. As Figure 1c shown is a schematic diagram of an example of compressing a matrix using the above compression method. For a 16x16 sparse matrix, only a, b, c, and d are non-zero elements. After compression storage, only the values and coordinates of these elements need to be stored, thus saving storage space.
[0008] When directly performing matrix calculations using the above traditional compression method for sparse matrices, generally there are two methods:
[0009] 1. First, decompress the compressed matrix to restore it to the original matrix, and then perform matrix operations.
[0010] 2. Directly use the compressed matrix for matrix operations. This method has a complex circuit and a relatively high circuit cost.
[0011] However, for the first method, it can only save storage space and cannot utilize the characteristics of the sparse matrix to accelerate matrix operations; for the second method, its circuit is complex and the cost is relatively high. Summary of the Invention
[0012] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the subsequent Detailed Implementation section. This Summary of the Invention section is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0013] To solve the above technical problems in the prior art, the embodiments of the present disclosure propose the following technical solutions:
[0014] In a first aspect, the embodiments of the present disclosure provide a matrix calculation circuit, including:
[0015] An instruction decoding circuit for decoding a matrix calculation instruction to obtain the starting address of a first matrix and the starting address of a second matrix, where the first matrix is the first compressed matrix of a data matrix, and the first matrix includes first data and the position data of the first data in the data matrix, and the first data is non-zero data in the data matrix;
[0016] A first data reading circuit for converting the data matrix into a bitmap matrix according to the position data, where the bitmap data in the bitmap matrix corresponds one-to-one with the data in the data matrix and is used to represent the position of the data;
[0017] A position data conversion circuit for converting the position data into bitmap data;
[0018] A control signal generation circuit for generating a second data reading control signal according to the bitmap data;
[0019] A second data reading circuit for generating a second data reading address according to the starting address of the second matrix and the second data reading control signal; and reading the second data according to the second data reading address;
[0020] A calculation circuit for calculating third data according to the first data and the second data.
[0021] Further, the position data conversion circuit includes:
[0022] A data cache circuit and a cache control circuit; where,
[0023] The cache control circuit is used to generate the storage address of the bitmap data according to the position data; and write a preset value into the data cache circuit according to the bitmap data storage address.
[0024] Further, the position data is the row coordinate and column coordinate of the first data in the data matrix, and generating the bitmap data storage address according to the position data includes:
[0025] Generating the storage address of the bitmap data according to the row coordinate, column coordinate, and the starting storage address of the data cache circuit.
[0026] Further, the instruction decoding circuit is further used to decode a matrix instruction to obtain the starting address of a third matrix;
[0027] The control signal generation circuit is further used to generate a third data storage control signal according to the bitmap data;
[0028] The matrix calculation circuit further includes:
[0029] A storage address generation circuit for generating a third data storage address based on the starting address of the third matrix and the third data storage control signal.
[0030] Further, the matrix calculation circuit further includes:
[0031] A first memory, a second memory, and a third memory;
[0032] Wherein the first memory is used to store the first data and the position data; releasing the first data to the calculation circuit and releasing the position data to the position data conversion circuit according to the read address of the first data;
[0033] The second memory is used to store the second data; releasing the second data to the calculation circuit according to the read address of the second data;
[0034] The third memory is used to save the third data to the storage location indicated by the storage address according to the third data storage address.
[0035] Further, the control signal generation circuit is further used to generate a first data read control signal according to the bitmap data, and the first data read control signal is used to control the first data read circuit to read the next first data and the next position data.
[0036] Further, the bitmap data includes column information of the bitmap data in the bitmap matrix, wherein,
[0037] The second data read circuit is used to: generate the read address of the second data according to the starting address of the second matrix and the column information.
[0038] Further, the bitmap data includes row information of the bitmap data in the bitmap matrix, and the control signal generation circuit is further used to:
[0039] Determine whether a row in the first matrix is calculated completely according to the row information;
[0040] In response to the calculation being completed, send an output instruction to the calculation circuit.
[0041] In a second aspect, an embodiment of the present disclosure provides a matrix calculation method, including:
[0042] Decoding a matrix calculation instruction to obtain the starting address of the first matrix and the starting address of the second matrix, wherein the first matrix is the first compression matrix of the data matrix, and the first matrix includes first data and the position data of the first data in the data matrix, and the first data is the non-0 data in the data matrix;
[0043] Generate a first data reading address based on the starting address of the first matrix;
[0044] Read the first data and the position data according to the first data reading address;
[0045] Convert the position data into bitmap data;
[0046] Generate a second data reading control signal according to the bitmap data;
[0047] Generate a second data reading address according to the starting address of the second matrix and the second data reading control signal;
[0048] Read the second data according to the second data reading address;
[0049] Calculate a third data based on the first data and the second data.
[0050] In a third aspect, an embodiment provides a processing core, including the matrix calculation circuit according to any one of the first aspect.
[0051] In a fourth aspect, an embodiment of the present disclosure provides a chip, including one or more of the processing cores described in the third aspect.
[0052] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: a memory for storing computer-readable instructions; and one or more processors for running the computer-readable instructions, so that when the processor runs, it implements any one of the matrix calculation methods in the foregoing first aspect.
[0053] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute any one of the matrix calculation methods in the foregoing first aspect.
[0054] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, characterized in that: it includes computer instructions, and when the computer instructions are executed by a computing device, the computing device can execute any one of the matrix calculation methods in the foregoing first aspect.
[0055] In an eighth aspect, an embodiment of the present disclosure provides a computing device, characterized in that it includes one or more of the chips described in the fourth aspect.
[0056] Embodiments of the present disclosure disclose a matrix calculation circuit and method. The matrix calculation circuit includes: an instruction decoding circuit for decoding a matrix calculation instruction to obtain the starting address of a first matrix and the starting address of a second matrix; a first data reading circuit for reading the first data and position data according to a first data reading address; a position data conversion circuit for converting the position data into bitmap data; a control signal generation circuit for generating a second data reading control signal according to the bitmap data; a second data reading circuit for reading the second data according to the second data reading control signal; and a calculation circuit for calculating a third data according to the first data and the second data. The above matrix calculation circuit converts position data into bitmap data through the position data conversion circuit, solving the technical problems of inability to accelerate operations or high circuit complexity and cost in the prior art.
[0057] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. In order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Combined with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.
[0059] Figures 1a - 1c Schematic diagram of the prior art of the present disclosure;
[0060] Figure 2 Schematic diagram of the structure of the matrix calculation circuit provided by the embodiment of the present disclosure;
[0061] Figure 3a Schematic diagram of compressing a data matrix using a bitmap matrix provided by the embodiment of the present disclosure;
[0062] Figure 3b Schematic diagram of an example of compressing a data matrix using a bitmap matrix provided by the embodiment of the present disclosure;
[0063] Figure 4a Schematic diagram of the structure of the position data conversion circuit provided by the embodiment of the present disclosure;
[0064] Figure 4b Schematic diagram of the position data conversion circuit converting position data into bitmap data provided by the embodiment of the present disclosure;
[0065] Figure 5Schematic diagram of compressing a sparse matrix using a traditional compression method provided by an embodiment of the present disclosure;
[0066] Figure 6 Schematic diagram of the storage format of a second matrix provided by an embodiment of the present disclosure;
[0067] Figure 7 Schematic flowchart of a matrix calculation method provided by an embodiment of the present disclosure;
[0068] Figure 8a Schematic diagram of an application example of the present disclosure;
[0069] Figure 8b Schematic diagram of compressing a data matrix using a bitmap matrix in the application of the present disclosure;
[0070] Figure 8c Schematic diagram of compressing a data matrix using a traditional compression method in the application of the present disclosure
[0071] Figure 8d Schematic diagram of the storage format of a second matrix in an application example of the present disclosure. Detailed implementation manners
[0072] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0073] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0074] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0075] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0076] It should be noted that the modification of "one" and "plural" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0077] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0078] Figure 2 It is a schematic diagram of the matrix calculation circuit provided for the embodiments of this disclosure. The matrix calculation circuit 200 provided in this embodiment includes:
[0079] An instruction decoding circuit ID (Instruction Decoder) 201, which is used to decode a matrix calculation instruction to obtain the starting address P_D of the first matrix and the starting address P_W of the second matrix; wherein the first matrix is the first compressed matrix of the data matrix, and the first matrix includes first data and the position data of the first data in the data matrix, and the first data is non-zero data in the data matrix;
[0080] A first data reading circuit (ADI_G) 202, which is used to generate a reading address A_Din of the first data according to the starting address P_D of the first matrix; read the first data D_data and the position data D_pos according to the reading address A_Din of the first data;
[0081] A position data conversion circuit (I2B) 203, which is used to convert the data matrix into a bitmap matrix according to the position data D_pos, and the bitmap data D_map in the bitmap matrix corresponds to the data in the data matrix one by one and is used to represent the position of the data;
[0082] A control signal generation circuit (map_P) 204, which is used to generate a second data reading control signal C_AW according to the bitmap data D_map;
[0083] A second data reading circuit (AW_G) 205, which is used to generate a reading address A_Win of the second data according to the starting address P_W of the second matrix and the second data reading control signal C_AW; read the second data Win according to the reading address A_Win of the second data;
[0084] A calculation circuit (EX) 206, which is used to calculate a third data Dout according to the first data D_data and the second data Win.
[0085] Exemplarily, the matrix calculation instruction is a matrix multiplication calculation instruction, which includes an instruction type, a storage start address of a first matrix participating in the instruction multiplication calculation, and a storage start address of a second matrix; in one embodiment, the first data in the first matrix is non-zero data in a data matrix in neural network convolution calculation, and the second matrix is a parameter matrix in neural network convolution calculation; optionally, the first matrix and / or the second matrix is a compressed matrix, where the compressed matrix is obtained by compressing an original matrix and only storing the non-zero data in the original matrix and the positions of the non-zero data in the original matrix. Optionally, the data in the compressed matrix is stored in the order of the data in the original matrix, first row by row and then column by column, and each row stores a first data and the position data of the first data.
[0086] Optionally, the bitmap data is data in a bitmap matrix, and the bitmap matrix is a matrix with the same dimension as the data matrix. Among them, the non-zero bitmap data in the bitmap matrix M1_map is used to indicate that the data at the position corresponding to the non-zero bitmap data in the data matrix is non-zero data, that is, the non-zero bitmap data in the bitmap matrix is used to indicate the non-zero data in the data matrix. For example, if the non-zero bitmap data in the bitmap matrix is represented by the numerical value 1, then the bitmap data with the numerical value 1 in the bitmap matrix represents that the data at the position corresponding to the numerical value 1 in the data matrix is non-zero data, and the bitmap data with the numerical value 0 in the bitmap matrix represents that the data at the position corresponding to the numerical value 0 in the data matrix is 0.
[0087] Optionally, the matrix calculation instruction further includes the dimensions of the data matrix represented by the first matrix and the parameter matrix represented by the second matrix participating in the instruction multiplication calculation, such as the number of rows and the width, and the start address of a third matrix, where the third matrix is the result matrix after performing the calculation on the first matrix and the second matrix. It can be understood that the storage start address of the matrix, the number of rows and columns of the matrix, and other parameters in the matrix calculation instruction can be represented in the form of register addresses, and the instruction decoding circuit obtains the corresponding data from the corresponding register addresses.
[0088] Figure 3a is a schematic diagram of the generation of the bitmap matrix. When the data matrix needs to be compressed, it can be compressed in the form of a one-dimensional matrix composed of non-zero data in the data matrix combined with the bitmap matrix, where a ij represents an element in the data matrix, where i ∈ (1, M), j ∈ (1, K); the one-dimensional matrix composed of non-zero data in the data matrix is in the order of a ijThe subscripts of [[]] are stored as a one-dimensional array in the order of rows first and columns second, i.e., row by row, and they are logically stored continuously. The bitmap data in the bitmap matrix corresponds one-to-one with the data in the data matrix, and is used to record the non-zero data in the data matrix. Optionally, it has the same dimension as the data matrix in form, i.e., the same number of rows and the same number of columns. The non-zero values in the bitmap matrix are used to represent the non-zero data in the data matrix. As Figure 3a shown, in the bitmap matrix, m represents the bitmap data in the bitmap matrix. Then when a ij ≠0, m ij ≠0. Since m ij is just an identification data, it can be represented by a small value, such as 1. At this time, each bitmap data in the bitmap matrix only requires 1 bit of storage space, and the size of the bitmap matrix is much smaller than that of the data matrix. Thus, by forming a one-dimensional matrix from the non-zero data in the data matrix to record the non-zero data in the data matrix, and using the bitmap matrix to record the positions of the non-zero data, the data matrix is compressed into two small matrices. It can be understood that in actual use, any other value can also be used in the bitmap matrix to represent the positions of the non-zero data in the data matrix, which will not be elaborated here.
[0089] In actual storage, all data is stored in units of bytes (Byte). For the bitmap matrix M1_map, the bitmap data in it is stored in bits (bit). 1 Byte can store 8 bits, that is, one byte can store 8 bitmap data of M1_map.
[0090] As Figure 3b shown is an example of generating a bitmap matrix. As Figure 3b shown, the data matrix M1_O includes four non-zero data a, b, c, d. The data matrix M1_O is compressed into a one-dimensional matrix M1 including 4 bits and a bitmap matrix M1_map. The bitmap data corresponding to the positions of the non-zero data a, b, c, d in the data matrix in the bitmap matrix M1_map is set to 1, and other positions are set to 0.
[0091] In the solution of the above embodiments of the present disclosure, the first matrix is a matrix obtained by compressing the data matrix using a traditional compression method. When actually calculating, the matrix calculation circuit uses the one-dimensional matrix and the bitmap matrix obtained by the compression method shown in 3a and 3b above. Therefore, the above position data conversion circuit I2B is required to convert the position data compressed by the traditional compression method into the bitmap data in the bitmap matrix. In this way, even if the data matrix is compressed using the traditional compression method, the matrix calculation can be performed using the method of combining the bitmap matrix and the one-dimensional matrix.
[0092] Figure 4aSchematic diagram of the position data conversion circuit in the embodiments of the present disclosure. As shown in FIG. 4, the position data conversion circuit 203 includes:
[0093] A data buffer circuit (DB) 401 and a buffer control circuit (Ctrl) 402; wherein,
[0094] The buffer control circuit 402 is configured to generate a storage address of bitmap data according to the position data; and write a preset value into the data buffer circuit 401 according to the storage address of the bitmap data.
[0095] Optionally, the position data is the row coordinate and column coordinate of the first data in the data matrix, and the generating the storage address of the bitmap data according to the position data includes:
[0096] Generating a storage address of the bitmap data according to the row coordinate, the column coordinate, and the storage start address of the data buffer circuit.
[0097] Exemplarily, the data buffer circuit 401 includes a plurality of storage locations, each storage location corresponding to 1-bit storage data, and the storage addresses of the plurality of storage locations are consecutive. After obtaining the row coordinate, the column coordinate, and the storage start address of the data buffer circuit, the storage address of the position represented by the row coordinate and the column coordinate in the data buffer circuit can be calculated. Let the storage start address of the data buffer circuit be DB_A, the row coordinate be X, the column coordinate be Y, the data matrix be a matrix of M*N, and the data in the data storage circuit be stored in row-major order. Then the storage location of the bitmap data corresponding to the position data is: A = DB_A + M*X + Y; if the storage data in the data storage circuit is stored in column-major order, the storage location of the bitmap data corresponding to the position data is: A = DB_A + X + N*Y. Thus, after receiving all the position data, the bitmap matrix is stored in the data buffer circuit.
[0098] Exemplarily, the data cache circuit 401 includes a plurality of storage locations, each storage location corresponding to 1-bit stored data. Each storage address of the data cache circuit corresponds to a logical address, and the logical address is the coordinate of the bitmap data in the bitmap matrix. For example, if the data cache circuit includes 100 storage locations, a logical address is assigned to each storage location in advance, so that the data cache circuit can represent a bitmap matrix smaller than 10*10. The correspondence between the physical address and the logical address of the data cache circuit is stored in the cache control circuit, and the data in each storage location of the data cache circuit is initialized to 0. When the cache control circuit receives the row coordinate and the column coordinate, it obtains the physical address of the bitmap data in the data cache circuit by querying the correspondence between the physical address and the logical address, and then sets the data corresponding to this physical address to 1. Thus, after receiving all the position data, the bitmap matrix is stored in the data cache circuit.
[0099] Figure 4b It is a schematic diagram of matrix conversion. As Figure 4b shown, the data matrix M1_O is a sparse matrix, which is compressed into a matrix M1_2 according to the traditional compression method, including the first data Data and the position data D_pos of the first data. The position data D_pos is converted into the bitmap data D_map in the bitmap matrix M1_map through the position data conversion circuit 203, and the Data data is delivered to the calculation circuit 206.
[0100] In an embodiment of the present disclosure, the first data reading circuit 202 receives the starting address P_D of the first matrix decoded by the instruction decoding circuit 201, generates the reading address A_Din of the first data according to the starting address P_D of the first matrix, and reads the first data D_Data and the position data D_pos according to the reading address A_Din of the first data. For example, read as Figure 5 shown a 11 and a 11 at the coordinates X = 1, Y = 1 in the data matrix. After each data fetch by the first data reading circuit 202, the starting address is updated, that is, after each fetch, the address of the first data fetched this time is used as the starting address for the next data fetch. If each first data and its position in the data matrix altogether occupy N storage locations, then each time P_D n = P_D n-1 + N, so that new first data and position data can be continuously fetched during the calculation.
[0101] Optionally, the first data is originally compressed in the compression mode of combining a bitmap matrix and a one-dimensional matrix. In this case, the first data reading circuit 202 receives the starting address P_D of the one-dimensional matrix and the starting address P_D_map of the bitmap matrix decoded by the instruction decoding circuit 201, generates the reading address A_Din of the first data according to the starting address P_D of the one-dimensional matrix, and generates the reading address A_map of the bitmap data according to the starting address P_D_map of the bitmap matrix; reads the first data D_Data according to the reading address A_Din of the first data, and reads the bitmap data D_map according to the reading address A_map of the bitmap data. As Figure 3a shown, in one example, if it is the first read, the first non-zero data a is read 11 and the 8 bitmap data m of the first byte in the bitmap data 11 to m 18 (assuming that the number of columns of the data matrix is at least 8). After each data fetch by the first data reading circuit 202, the starting address is updated, that is, after each fetch, the address of the first data fetched this time plus the length of the data is used as the starting address for the next data fetch, and the address of the bitmap data fetched this time plus the length of the data is used as the starting address for the next data fetch; if each first data occupies N storage positions, then each time P_D n = P_D n-1 + N, so that new first data can be continuously fetched in the calculation; each bitmap data only occupies 1 bit, so 8 bitmap data are fetched at a time when fetching bitmap data, so the data fetching frequency of the bitmap data is different from that of the first data. Each time the starting address of the bitmap matrix is updated, only 1 needs to be added, that is, P_D_map n = P_D_map n-1 + 1.
[0102] In the embodiment of the present disclosure, the control signal generation circuit (map_P) 204 generates a second data reading control signal (C_AW) according to the bitmap data D_map. Optionally, the second data reading control signal (C_AW) includes the position of the bitmap data corresponding to the first data in the bitmap data.
[0103] In the embodiment of the present disclosure, the second data reading circuit 205 generates the reading address A_Win of the second data according to the starting address P_W of the second matrix and the position of the bitmap data; reads the second data Win according to the reading address A_Win of the second data. Optionally, the second data reading circuit 205 generates the reading address A_Win of the second data according to the starting address P_W of the second matrix and the number of columns of the bitmap data, where the number of columns of the bitmap data indicates which column the bitmap data is in the bitmap matrix.
[0104] In one embodiment, all elements in the second matrix are non-zero elements. At this time, the second matrix is not compressed and is directly stored in the original form of the parameter matrix; its stored logical form is as Figure 6 shown, where P_W is the storage address of the first element b 11 of the second matrix. Logically, it is stored continuously in the form of a matrix, while physically, it can be stored continuously or discontinuously, which is not limited in the present disclosure. Through the starting address P_W and the number of columns of the bitmap data, the row address of the second matrix can be generated, and a row of second data corresponding to the number of columns of the bitmap data in the second matrix can be read out for subsequent matrix calculations. When the first data taken out changes, the bitmap data changes accordingly, so the number of columns of the bitmap data also changes, and the second data read out will continuously change with the number of columns of the bitmap data, thereby changing the second data used in each calculation. Optionally, if the second matrix is a sparse matrix, the second matrix can be compressed in any one of the two ways of compressing the data matrix into the first matrix.
[0105] In the present disclosure, the computing circuit 206 calculates a third data based on the first data and the second data. Optionally, if the matrix calculation instruction is a matrix multiplication instruction, then when the matrix multiplication instruction is executed by the computing circuit, it is split into a multiplication operation of the first data and the second data and an accumulation calculation of the third data, that is, the computing circuit finally completes the multiply-accumulate calculation. In each clock cycle, the computing circuit multiplies the first data and the second data obtained in the current clock cycle to obtain a multiplication calculation result, and adds the multiplication calculation result to the third data obtained in the previous clock cycle to obtain the third data in the current clock cycle. And so on, the computing circuit 206 continuously performs the multiply-accumulate calculation until the calculation is completed.
[0106] Optionally, the computing circuit 206 includes a plurality of computing units. Exemplarily, the number of columns of the computing unit is the same as the number of columns of the second matrix. That is, when the second data reading circuit 206 reads out a row of second data at a time according to the position of the bitmap data, the row of second data can be calculated in parallel with the first data respectively.
[0107] Optionally, the control signal generation circuit 204 is further configured to generate a first data reading circuit control signal (C_ADI) according to the bitmap data. The reading control signal of the first data is used to control the first data reading circuit 202 to read the next first data and / or control the first data reading circuit 202 to read the next bitmap data (determined according to the compression form of the first matrix. If it is compressed according to the traditional compression method, only the first data needs to be read, so the first data and its position data are stored together; if it is compressed according to the compression method combining the bitmap matrix and the one-dimensional matrix, the next first data and the next bitmap data need to be read).
[0108] Among them, the bitmap data includes the value of the bitmap data and the position of the bitmap data. The position of the bitmap data represents the position of the bitmap data in the bitmap matrix, that is, the position of the first data in the data matrix; optionally, the position of the bitmap data is an indirect position, that is, the read position of the bitmap data is the logical position of the bitmap data, and then the position of the bitmap data in the bitmap matrix is calculated through the number of rows H_map and the number of columns W_map of the bitmap matrix and the storage position of the bitmap data; for example, assume that the number of rows of the bitmap matrix is 4 and the number of columns is 4, then the position of the 11th bitmap data in the logical storage position in the bitmap matrix is the 3rd row and the 3rd column; thus, the control signal generation circuit only needs to determine the logical storage position of the currently processed bitmap data to determine the position of the bitmap data in the bitmap matrix. In this optional embodiment, since the position of the bitmap data does not need to be stored, more storage space is saved.
[0109] Optionally, the position of the bitmap data is directly recorded in the bitmap data, that is, the bitmap data includes the value of the bitmap data and the position (h, w) of the bitmap data in the bitmap matrix, where h represents the row number of the bitmap data in the bitmap matrix, and w represents the column number of the bitmap data in the bitmap matrix. In this optional embodiment, since the position of the bitmap data is already included in the bitmap data, it requires additional storage space, but it can save the calculation amount of calculating the position of the bitmap data. For example, when the next first data needs to be fetched, the bitmap data is fetched first. When the value of the bitmap data is 1, it means that the first data corresponding to the bitmap data is a non-zero data, and the control signal generation circuit 204 generates the first data reading circuit control signal to control the first data reading circuit to read the first data corresponding to the bitmap data; when the bitmap data is used up, for example, 1 byte of bitmap data is read at a time, and when all 8 bitmap data are traversed, the control signal generation circuit 204 generates the first data reading circuit control signal to control the first data reading circuit to read the next bitmap data.
[0110] It can be understood that, in this example, the next bitmap data is 8-bit bitmap data of one byte. It can be understood that the above method of reading the first data is only an example and does not constitute a limitation to the present disclosure; for example, the control signal generation circuit may also generate a control signal according to a periodic signal to control the first data reading circuit to read the next first data in each period; or read the next first data according to the feedback signal of the calculation circuit, that is, whenever the calculation circuit finishes calculating the current first data, it sends a feedback signal to the control signal generation circuit to cause the first data to read the next first data. Other methods of reading the first data will not be elaborated.
[0111] Optionally, the position of the bitmap data further includes the row number of the bitmap data in the bitmap matrix, where the row number indicates which row the bitmap data is in the bitmap matrix. The control signal generation circuit 204 is further configured to: determine whether a row in the first matrix is calculated based on the row number of the bitmap data in the bitmap matrix; in response to the calculation being completed, send an output instruction (C_RE) to the calculation circuit 206.
[0112] As Figure 2 shown, the control signal generation circuit 204 receives the bitmap data and determines whether a row in the first matrix is calculated by the row number of the bitmap data. Since the dimensions of the data matrix and the bitmap matrix are the same, a row in the bitmap matrix corresponds to a row in the data matrix. Therefore, when the row number of the bitmap data changes, it indicates that a row in the data matrix is calculated. Therefore, the row number in the position of the bitmap data can be used to determine whether a row of the first data in the first matrix is calculated. Exemplarily, the row number of the current bitmap data can be determined by the logical address of the current bitmap data and the number of columns of the bitmap matrix. Let the logical address of the current bitmap data be add, which represents the position of the current bitmap data among all the bitmap data in the bitmap matrix sorted by row first and column second. Let the number of columns of the bitmap matrix be W_map, then the row number of the current bitmap data For example, for a 4*4 bitmap data, where in the order of row first and column second, the row number where the 5th bitmap data is located is i.e., the second row.
[0113] Specifically, the judgment steps are as follows:
[0114] Compare whether the row number of the bitmap data used this time is the same as the row number of the bitmap data used last time;
[0115] If they are the same, then a row of the first data in the first matrix has not been calculated yet; or,
[0116] If they are different, then a row of the first data in the first matrix is calculated.
[0117] That is, compare whether the X coordinate of the currently read bitmap data is the same as the X coordinate of the previously read bitmap data. If they are the same, it means that a row of data has not been completely read. If they are different, it means that a new line of bitmap data has been read, indicating that all the first data in a row have participated in the calculation. In response to the completion of the calculation, that is, the X coordinate of the currently read bitmap data is different from the X coordinate of the previously read bitmap data, an output instruction C_RE is sent to the calculation circuit 206 to enable the calculation circuit 206 to output the calculated third data of a row.
[0118] Optionally, the control signal generation circuit 204 can also obtain the width and height of the bitmap matrix through the instruction decoding circuit. Since the bitmap matrix is stored row by row first and then column by column, the number of rows and columns of the bitmap data can be obtained therefrom.
[0119] Optionally, the instruction decoding circuit 201 is further configured to decode the matrix calculation instruction to obtain the starting address of the third matrix. The control signal generation circuit 204 is further configured to generate a third data storage control signal (C_ADO) according to the bitmap data. The matrix calculation circuit 200 further includes: a storage address generation circuit (ADO_G) 207, which generates the storage address A_Dout of the third data according to the starting address P_O of the third matrix and the third data storage control signal. Wherein, the third matrix is a matrix obtained by performing calculations on the first matrix and the second matrix. Exemplarily, the position of the bitmap data in the bitmap matrix in this embodiment is the row number of the bitmap data in the bitmap matrix. When the matrix calculation is matrix multiplication calculation, the calculation circuit can obtain a row of third data in the third matrix for each row of first data calculated in the first matrix, that is, each calculation unit in the calculation circuit 206 outputs one of the third data in a row of the third matrix. Thus, the row number of the bitmap data in the bitmap matrix can determine which row of data in the third matrix the output row of third data is. As Figure 2 shown, the storage address generation circuit 207 receives the third data storage control signal to determine which row of the third matrix is output, and then determines the storage address A_Dout of the output row of third data according to the starting address P_O of the third matrix and the row number.
[0120] Optionally, after receiving the output instruction C_RE sent by the control signal generation circuit 204, the storage address generation circuit 207 sends out the storage address A_Dout of the third data.
[0121] Optionally, as Figure 2 shown, the matrix calculation circuit 200 further includes: a first memory 208, a second memory 209, and a third memory 210; wherein,
[0122] The first memory 208 is used to store the first data and the position data of the first data or the first data and the bitmap data (as described above, determined by the compression method of the data matrix); release the first data corresponding to the read address to the computing circuit 206 according to the read address of the first data, release the position data of the first data to the position data conversion circuit 203, or release the bitmap data to the position data conversion circuit 203 according to the read address of the bitmap data;
[0123] The second memory 209 is used to store the second data: release the second data to the computing circuit 206 according to the read address of the second data;
[0124] The third memory 210 is used to save the third data to the storage location indicated by the storage address according to the storage address of the third data.
[0125] In the embodiment of the present disclosure, the first matrix, that is, the compression matrix of the data matrix, is stored in the first memory 208, which includes the first data and the position data of the first data; or the first data and the bitmap matrix of the first data; the second matrix, that is, the parameter matrix or the compression matrix of the parameter matrix, is stored in the second memory 209; the third matrix, that is, the output matrix, is stored in the third memory 210. The first memory 208, the second memory 209, and the third memory 210 can be accessed independently and in parallel, so that data fetching, parameter fetching, calculation, and result storage can be performed in parallel, thereby improving the efficiency of matrix calculation.
[0126] It can be understood that the first data reading circuit 202 and the first memory 208 are connected by an address bus and a data bus, where the address bus is used to transmit addresses. The first data reading circuit enables the first memory 208 to release data to the data bus through the address circuit, and the circuit connected to the data bus can obtain the required first data and the position of the first data in the first matrix; similarly, the connection relationship between the second data reading circuit 205 and the second memory 209 is similar and will not be described in detail. The address bus of the third memory 210 is connected to the storage address generation circuit 207 for transmitting the storage address of the third data generated by the storage address generation circuit 207, and the data bus of the third memory 210 is connected to the computing circuit 206 for transmitting the third data calculated by the computing circuit 206.
[0127] It can be understood that in the above specific embodiments, the position of the bitmap data in the bitmap matrix (i.e., the corresponding position in the data matrix) is used to determine the second data to be read. In the actual implementation process, the first matrix and the second matrix can also be inverted, that is, the first data to be read is determined according to the position of the bitmap data in the bitmap matrix corresponding to the second matrix. That is, for matrix multiplication calculation, each time a second data is read, a corresponding column of first data needs to be read from the first matrix, and then the second data is read column by column. In this way, each time a column of second data is calculated, a column of third data is obtained. When outputting the third data, according to the number of columns of the bitmap data in the bitmap matrix of the second data, it can be determined which column of data in the third matrix the output third data is, and a column of third data is stored according to the starting address of the third data and the number of columns of the bitmap data. Other situations can be obtained according to the above process and the specific matrix calculation process, which will not be elaborated here.
[0128] Figure 7 The flowchart of the matrix calculation method provided by the embodiments of the present disclosure is as follows. As Figure 7 shown, the method includes the following steps:
[0129] Step S701: Decode the matrix calculation instruction to obtain the starting address of the first matrix and the starting address of the second matrix, where the first matrix is the first compression matrix of the data matrix, and the first matrix includes the first data and the position data of the first data in the data matrix, and the first data is the non-0 data in the data matrix;
[0130] Step S702: Generate a first data reading address according to the starting address of the first matrix;
[0131] Step S703: Read the first data and the position data according to the first data reading address;
[0132] Step S704: Convert the data matrix into a bitmap matrix according to the position data, and the bitmap data in the bitmap matrix corresponds to the data in the data matrix one by one, and is used to represent the position of the data;
[0133] Step S705: Generate a second data reading control signal according to the bitmap data;
[0134] Step S706: Generate a second data reading address according to the starting address of the second matrix and the second data reading control signal;
[0135] Step S707: Read the second data according to the second data reading address;
[0136] Step S708: Calculate the third data according to the first data and the second data.
[0137] Furthermore, the step S704 further includes:
[0138] generating a storage address of the bitmap data according to the position data;
[0139] A preset value is written into a data cache circuit according to a storage address of the bitmap data.
[0140] Further, generating a storage address of the bitmap data according to the position data includes:
[0141] The storage address of the bitmap data is generated according to the row coordinates, the column coordinates and the storage head address of the data cache circuit.
[0142] Furthermore, the method further comprises:
[0143] Decode the matrix instruction to obtain the first address of the third matrix;
[0144] generating a third data storage control signal according to the bitmap data;
[0145] A third data storage address is generated according to the first address of the third matrix and the third data storage control signal.
[0146] Furthermore, the method further comprises:
[0147] releasing the first data to the calculation circuit according to the read address of the first data, and releasing the position data to the position data conversion circuit;
[0148] releasing the second data to the computing circuit according to the second data reading address;
[0149] The third data is stored in the storage location indicated by the storage address.
[0150] Furthermore, the method further comprises:
[0151] A first data read control signal is generated according to the bitmap data, and the first data read control signal is used to control the first data read circuit to read the next first data and the next position data.
[0152] Furthermore, the method further comprises:
[0153] A read address for the second data is generated according to the first address of the second matrix and the column information.
[0154] Furthermore, the method further comprises:
[0155] Determine whether a row in the first matrix has been calculated according to the row information;
[0156] In response to the completion of the calculation, an output instruction is sent.
[0157] In the foregoing, although the steps in the above method embodiments are described in the above order, those skilled in the art should understand that the steps in the embodiments of the present disclosure do not necessarily need to be executed in the above order, and they can also be executed in a reverse order, in parallel, in a cross order, etc. Moreover, based on the above steps, those skilled in the art can also add other steps. These obvious variations or equivalent replacement methods should also be included in the protection scope of the present disclosure and will not be elaborated herein.
[0158] The embodiments of the present disclosure further provide a processing core, and the processing core includes any one of the matrix calculation circuits in the above embodiments.
[0159] The embodiments of the present disclosure further provide a chip, and the chip includes one or more of the above processing cores.
[0160] The working process of the matrix calculation circuit in the embodiments of the present disclosure is described below with a practical application scenario. Figure 8a In this example application scenario, the data matrix M1_O is multiplied by the parameter matrix M2 to obtain the output matrix M3. Where M1_O is a 4*4 sparse data matrix, which includes a large number of 0 elements, and M2 is a 4*2 parameter matrix, and the two are multiplied to obtain a 4*2 output matrix.
[0161] In this example, all matrix elements are FP32, that is, 32 bits represent one data, so each element in each matrix requires 32 / 8 = 4B of storage space. The parameter matrix includes 2 columns of elements, so the calculation circuit 206 can be set as a vector processing unit composed of two multiply-accumulators.
[0162] Such as Figure 8bShown are two matrices obtained after compressing the sparse data matrix M1_O according to the compression method of the bitmap matrix, namely, the first matrix M1_1 stored in the first memory and the bitmap matrix M1_1_map; wherein, the first matrix M1_1 includes the non-zero data 1, 2, 3, 4 in the data matrix M1_O, and is stored in sequence (i.e., row by row) in the order of first row and then column. Each element in the first matrix M1_1 requires 4B to store. After the first data reading circuit 202 finishes reading each first data, the address P_D is automatically incremented by 4B to obtain the address of the next first data, i.e., P_D = P_D + 4. Since the data in M1_map are all bit data and there are only 4 data in each row, when storing, one row of M1_1_map only needs to occupy 4bit of storage space. In this way, 2 rows of M1_1_map can be stored in one byte, thus saving storage space. After every 2 rows are calculated, the fetch address of M1_1_map changes to: P_D_map = P_D_map + 1.
[0163] As Figure 8c Shown is the compressed matrix M1_2 obtained after compressing the sparse data matrix M1_O according to the traditional compression method. It includes the non-zero data in the data matrix M1_O and the position coordinates X and Y of the non-zero data in the data matrix M1_O. Since the values of the position coordinates of the matrix are generally relatively small, in this example, 2B of storage space can be used to store the position of the first data in the data matrix, with X occupying 1B and Y occupying 1B. Thus, each element in the matrix M1_2 requires 4B + 1B + 1B = 6B to store. After the first data reading circuit 202 finishes reading each first data position data, the address P_D is automatically incremented by 6B to obtain the address of the next first data, i.e., P_D = P_D + 6.
[0164] As Figure 8d Shown is the logical storage format of the second matrix, i.e., the parameter matrix, which is stored in row order. The size of the second data in each row is 2 * 4 = 8B. The storage start address of each row is determined by the initial address P_W of the second matrix and the row number N_r, i.e., P_W + 8 * N_r. Each time a row of second data is read according to P_W, and then the value of P_W is updated.
[0165] When performing calculations, the instruction decoding circuit 201 decodes the start address P_D_1 of M1_1, the start address P_map of M1_map, or the start address P_D_2 of M1_2 (determined by the type of the stored compressed matrix), and the start address P_W of M2.
[0166] The following takes decoding the start address P_D_2 of M1_2 and the start address P_W of M2 as an example for illustration. Decoding the start address P_D_1 of M1_1 and the start address P_map of M1_map is similar to this case, except that it is not necessary to convert the position data into bitmap data. After obtaining the bitmap data, the calculation methods of the two are the same.
[0167] After obtaining the start address P_D_2 of M1_2 and the start address P_W of M2 as described above, the first data reading circuit 202 reads the first first data 1 and the position data (1, 1) of the first data 1 from the first memory 208 according to the start address P_D_1, and inputs the position data (1, 1) into the position data conversion circuit to restore it into bitmap data. In the bitmap data, the bitmap data with a value of 1 indicates that there is non-zero data at the position corresponding to the bitmap data in the data matrix, that is, the first data 1; the control signal generation circuit 204, after obtaining the bitmap data of the first data 1, can obtain that the position of the bitmap data in the bitmap matrix is (1, 1), that is, the first row and the first column. From this, the control signal of the second data reading circuit 205 can be generated, and the column number 1 of the bitmap data is carried in the control signal; the second data reading circuit 205 then reads the first row second data 1 and 2 corresponding to the column number 0 from the second memory 209 according to the start address P_W of the second matrix and the column number 1 of the bitmap data; when obtaining the bitmap data of the first data 2, it can be obtained that the position of the bitmap data in the bitmap matrix is (2, 3), that is, the second row and the third column. From this, the control signal of the second data reading circuit 205 can be generated, and the column number 3 of the bitmap data is carried in the control signal; the second data reading circuit 205 then reads the third row second data 1 and 2 corresponding to the column number 3 from the second memory 209 according to the start address P_W of the second matrix and the column number 3 of the bitmap data, that is, the column number of the bitmap data determines the reading address of the second data generated by the second data reading circuit 205. Let the column number of the bitmap data be Col, then the reading address of the second data is: P_W + 8*(Col - 1). Note that in this example, for the convenience of understanding, the column number is marked starting from 1. In fact, it can also be marked starting from 0, that is, the row and column numbers of the first data are (0, 0). At this time, the reading address of the second data needs to be adjusted to: P_W + 8*Col; in short, through the column number of the bitmap data, the second data reading circuit 205 can determine the start address of a row of second data, and then read out this row of second data.
[0168] The computing circuit 206 is composed of a plurality of computing units, and each computing unit includes a multiply-accumulator; after fetching the first data and a row of second data, the first data is transmitted to an input port of the two multiply-accumulators, and the two second data are respectively transmitted to the other input ports of the above two multiply-accumulations. The multiply-accumulator performs a multiply-accumulation calculation according to the received first data and second data, multiplies the first data and the second data, and accumulates the result obtained from the previous multiply-accumulation calculation to obtain the third data. The computing circuit 206 performs the multiply-accumulation calculation in the above manner until the number of columns of the bitmap data changes, indicating that the calculation of a row of the first data is completed.
[0169] Taking the above example, when the second first data 2 is read, the number of columns of the corresponding bitmap data is 3, which is different from the X coordinate 1 of the first first data 1, indicating that the first data corresponding to a row of data in the data matrix has been calculated. At this time, the control signal generation circuit 204 generates an output instruction C_RE and sends it to the computing circuit 206, and sends a storage control signal C_ADO to the storage address generation circuit 207; after receiving the storage control signal C_ADO, the storage address generation circuit 207 generates a storage address for a row of the third data of the output matrix M3, and the computing circuit 206 releases a row of the third data to the data bus. The third memory 210 stores a row of the third data on the data bus according to the storage address of the third data. Each time the calculation of a row of output is completed, the storage address generation circuit 207 calculates the storage address of the third data according to the number of rows where the bitmap data is located. Each row of the output matrix M3 has 2 FP32 data, so the storage size of each row is 2 * 4 = 8B. If the starting address of the output matrix storage is P_M and the number of rows where the bitmap data is located is Row, then the storage address of this output row is: P_M = P_M + 8 * Row.
[0170] Iteratively execute the above steps until all the first data are read out, and obtain the output matrix in the third memory, that is, the matrix calculation result of the first matrix and the second matrix.
[0171] The disclosed embodiment provides a processing core, including the matrix calculation circuit as described in the above embodiment.
[0172] The disclosed embodiment of the present disclosure provides a chip, including one or more of the above-mentioned processing cores.
[0173] The disclosed embodiment of the present disclosure provides an electronic device, including: a memory for storing computer-readable instructions; and one or more processors for running the computer-readable instructions, so that when the processor runs, it implements any one of the matrix calculation methods in the embodiments.
[0174] An embodiment of the present disclosure also provides a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute any one of the matrix calculation methods in the foregoing embodiments.
[0175] An embodiment of the present disclosure also provides a computer program product, characterized in that it includes computer instructions, and when the computer instructions are executed by a computing device, the computing device can execute any one of the matrix calculation methods in the foregoing embodiments.
[0176] An embodiment of the present disclosure also provides a computing device, characterized in that it includes any one of the chips in the foregoing embodiments.
[0177] The flowcharts and block diagrams in the accompanying drawings of the present disclosure illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a task, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0178] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0179] The functions described above herein can be at least partially performed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0180] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
Claims
1. A matrix calculation circuit, characterized in that Including: An instruction decoding circuit for decoding a matrix calculation instruction to obtain the starting address of a first matrix and the starting address of a second matrix, where the first matrix is the first compressed matrix of a data matrix, and the first matrix includes first data and position data of the first data in the data matrix, and the first data is non-zero data in the data matrix; A first data reading circuit for generating a first data reading address according to the starting address of the first matrix; Reading the first data and the position data according to the first data reading address; A position data conversion circuit for converting the data matrix into a bitmap matrix according to the position data, where the bitmap data in the bitmap matrix corresponds one-to-one with the data in the data matrix and is used to represent the position of the data; A control signal generating circuit for generating a second data reading control signal according to the bitmap data; A second data reading circuit for generating a second data reading address according to the starting address of the second matrix and the second data reading control signal; Reading the second data according to the second data reading address; A calculation circuit for calculating third data according to the first data and the second data; The instruction decoding circuit is further configured to decode a matrix instruction to obtain the starting address of a third matrix; The control signal generating circuit is further configured to generate a third data storage control signal according to the bitmap data; The matrix calculation circuit further includes: A storage address generating circuit for generating a third data storage address according to the starting address of the third matrix and the third data storage control signal.
2. The matrix calculation circuit according to claim 1, wherein The position data conversion circuit includes: A data cache circuit and a cache control circuit; where The cache control circuit is configured to generate a storage address of the bitmap data according to the position data; and write a preset value into the data cache circuit according to the storage address of the bitmap data.
3. The matrix calculation circuit according to claim 2, wherein The position data is the row coordinate and column coordinate of the first data in the data matrix, and generating the storage address of the bitmap data according to the position data includes: Generating the storage address of the bitmap data according to the row coordinate, the column coordinate, and the starting storage address of the data cache circuit.
4. The matrix calculation circuit according to claim 1, wherein The matrix calculation circuit further includes: A first memory, a second memory, and a third memory; Where the first memory is used to store the first data and the position data; releasing the first data to the calculation circuit according to the reading address of the first data, and releasing the position data to the position data conversion circuit; The second memory is used to store the second data; releasing the second data to the calculation circuit according to the second data reading address; The third memory is used to save the third data to the storage position indicated by the storage address according to the third data storage address.
5. The matrix calculation circuit according to claim 1, characterized in that: The control signal generation circuit is further configured to generate a first data read control signal according to the bitmap data, and the first data read control signal is used to control the first data read circuit to read the next first data and the next position data.
6. The matrix calculation circuit according to claim 1, wherein The bitmap data includes column information of the bitmap data in the bitmap matrix, where The second data read circuit is configured to: Generate a read address of the second data according to the start address of the second matrix and the column information.
7. The matrix calculation circuit according to any one of claims 1-6, characterized in that, The bitmap data includes row information of the bitmap data in the bitmap matrix, and the control signal generation circuit is further configured to: Determine whether a row in the first matrix is calculated completely according to the row information; in response to the completion of the calculation, send an output instruction to the calculation circuit.
8. A matrix calculation method, characterized in that, Comprising: Decoding a matrix calculation instruction to obtain a start address of a first matrix and a start address of a second matrix, where the first matrix is a first compression matrix of a data matrix, and the first matrix includes first data and position data of the first data in the data matrix, and the first data is non-zero data in the data matrix; Generating a first data read address according to the start address of the first matrix; Reading the first data and the position data according to the first data read address; Converting the data matrix into a bitmap matrix according to the position data, and the bitmap data in the bitmap matrix corresponds to the data in the data matrix one by one and is used to represent the position of the data; Generating a second data read control signal according to the bitmap data; Generating a second data read address according to the start address of the second matrix and the second data read control signal; Reading the second data according to the second data read address; Calculating a third data according to the first data and the second data; The method further includes: Decoding a matrix instruction to obtain a start address of a third matrix; Generating a third data storage control signal according to the bitmap data; Generating a third data storage address according to the start address of the third matrix and the third data storage control signal.
9. A processing core, characterized in that, Comprising the matrix calculation circuit according to any one of claims 1-7.
Citation Information
Patent Citations
Data processing method and device
CN109597647A
Matrix processing method and apparatus, and logic circuit
CN111010883A