Matrix calculation circuit, method, electronic device and computer-readable storage medium

By designing a matrix calculation circuit, using cached and reordered data reading circuits, efficient sparse matrix calculation is realized, solving the problems of low data utilization and complex number address acquisition, and improving the computing speed and chip computing power.

CN114168894BActive Publication Date: 2025-07-08STREAM COMPUTING INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010955584.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-11
Publication Date
2025-07-08
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

In the prior art, when performing matrix calculation, the data utilization rate is low and the number-choice address is complex, which affects performance performance. Especially when processing sparse matrices, it is impossible to effectively utilize the calculation unit and improve the computing speed.

Method used

A matrix calculation circuit is designed, including the first and second data reading circuits. By buffering and reordering the data and position information of the sparse matrix, a calculation unit array is used to process multiple data simultaneously to realize efficient matrix multiplication calculation.

Benefits of technology

It improves data utilization and computing speed, saves storage space and data bandwidth, improves the effective computing power of the chip, and can quickly process large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114168894B_ABST
    Figure CN114168894B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a matrix calculation circuit, method, electronic device, and computer-readable storage medium. The matrix calculation circuit includes: a first data reading circuit configured to read and cache first data in a first matrix and position information of the first data, where the first matrix is a compressed matrix of a data matrix; a second data reading circuit configured to read and cache second data in a second matrix according to the position information of the first data; and a calculation circuit configured to calculate third data based on the first data and the second data. The above matrix calculation circuit controls the reading of the second data through the position information of the first data read out, solving the technical problems in the prior art that only single data calculation can be performed during matrix calculation and the fetch address calculation is complex.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of processors, and particularly to a matrix calculation circuit, method, electronic device, and computer-readable storage medium. Background Art

[0002] With the development of science and technology, human society is rapidly entering the intelligent era. An important feature of the intelligent era is that people obtain more and more types of data, and the amount of data obtained is increasing, while the requirement for the speed of processing data is getting higher and higher. A chip is the cornerstone of task allocation, which fundamentally determines people's ability to process data. From the perspective of application fields, there are mainly two routes for chips: one is the general-purpose chip route, such as the CPU (central processing unit), etc., which can provide great flexibility, but the effective computing power is relatively low when processing algorithms in specific fields; the other is the dedicated chip route, such as the TPU (tensor processing unit), etc., which can exert higher effective computing power in certain specific fields, but when facing flexible and general fields, their processing ability is relatively poor or even unable to process. Due to the wide variety and large quantity of data in the intelligent era, it is required that the chip not only has extremely high flexibility to process algorithms in different fields and keep changing with each passing day, but also has extremely strong processing ability to quickly process a large and rapidly growing amount of data.

[0003] In neural network computing, convolution computing accounts for most of the total amount of operations, and convolution computing can be converted into matrix multiplication computing. Therefore, to improve the throughput, reduce the latency, and enhance the effective computing power of the chip in neural network tasks, the key lies in improving the speed of matrix multiplication computing.

[0004] Many matrices composed of data in neural networks (here the data includes parameter data and input data in neural networks) are sparse matrices, that is, there are a large number of elements in the matrix with a value of 0. In order to reduce the storage amount and bandwidth occupation of data in neural network computing, the sparse matrix will be compressed for storage; in order to improve the matrix operation speed, the sparse matrix operation will be optimized.

[0005] Figure 1a FIG. is a schematic diagram of matrix multiplication calculation in a neural network. As Figure 1a shown, M1 is a data matrix, M2 is a parameter matrix, and M is an output matrix. One row of data in M1 and one column of parameters in M2 are multiplied and added to obtain one data in M. Among them Figure 1a for the two matrices M1 and M2, one of them may be a sparse matrix, or both may be sparse matrices.

[0006] As Figure 1bThe following is a schematic diagram of matrix compression. For the storage of sparse matrices, a common compression method can be adopted: only store non-zero elements. When storing the values of these non-zero elements, the position information of the elements in the matrix will also be stored, that is, the relative coordinates X and Y of the elements in the matrix. Among them, X represents the row number of the matrix, and Y represents the column number of the matrix. This method takes data and coordinates as a data structure and stores them in units of this data structure. As Figure 1b shown, taking an MxN matrix as an example, it is compressed from the left MxN matrix to the right compressed matrix. Each data structure in the compressed matrix represents the non-zero data in the left matrix and the coordinates of this non-zero data in the matrix.

[0007] In a sparse matrix, since some elements in the matrix have a value of 0 and these 0 elements do not need to be stored, adopting this compression method can effectively reduce the storage capacity of the matrix. As Figure 1c shown is a schematic diagram of an example of compressing a matrix using the above compression method. For a 16x16 sparse matrix, only a, b, c, and d are non-zero elements. After compression storage, only the values and coordinates of these elements need to be stored, thus saving storage space.

[0008] When performing matrix operations of M1xM2, the compressed matrix is used as the matrix for actual data fetching. However, the above technical solution has the following disadvantages: 1. When performing matrix operations, the utilization rate of data is low, and usually only independent operation units can be used to calculate single data; 2. According to the data coordinates of the compressed matrix, calculating the data fetching address is complex, which affects the performance. Summary of the Invention

[0009] This Summary of the Invention section is provided to introduce concepts in a concise form, which will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0010] To solve the above technical problems in the prior art, the embodiments of the present disclosure propose the following technical solutions:

[0011] In a first aspect, the embodiments of the present disclosure provide a matrix calculation circuit, including:

[0012] A first data reading circuit, configured to read and cache the first data in the first matrix and the position information of the first data, where the first matrix is a compressed matrix of a data matrix;

[0013] A second data reading circuit, configured to read and cache the second data in the second matrix according to the position information of the first data;

[0014] A calculation circuit for calculating third data based on the first data and the second data.

[0015] Furthermore, the first data reading circuit further includes:

[0016] A first data cache circuit, a first data sorting circuit, and a first control circuit;

[0017] Wherein, the first control circuit is used to generate a first data reading address according to the starting address of the first matrix;

[0018] The first data cache circuit is used to cache the first data read out according to the first data reading address and the position information of the first data;

[0019] The first data sorting circuit is used to re - sort the first data position information and the first data in a one - to - one position - corresponding manner according to the first data position information in the first data cache circuit. Among them, in the re - sorting result, the data in the same row of the data matrix is still in the same row, and the data in the same column of the data matrix is still in the same column.

[0020] Furthermore, the first data sorting circuit is further used to:

[0021] Send the position information of the first data to the second data reading circuit.

[0022] Furthermore, the second data reading circuit further includes:

[0023] A second data cache circuit and a second control circuit;

[0024] Wherein, the second control circuit is used to generate a second data reading address according to the starting address of the second matrix and the position information of the first data;

[0025] The second data cache circuit is used to cache the second data read out according to the second data reading address.

[0026] Furthermore, the second control circuit is used to generate a second data reading address according to the starting address of the second matrix and the position information of the first data, including:

[0027] The second control circuit is used to generate the second data reading address according to the column information in the starting address of the second matrix and the position information of the first data.

[0028] Furthermore, the calculation circuit includes:

[0029] An array of calculation units, where the array of calculation units includes a plurality of calculation units;

[0030] The row computing units in the computing unit array simultaneously receive a row of the plurality of second data;

[0031] The column computing units in the computing unit array simultaneously receive a column of the plurality of first data.

[0032] Further, the computing circuit for calculating third data according to the first data and the second data includes:

[0033] The computing circuit receives a column of the first data output by the first data sorting circuit; receives a row of the second data output by the second data caching circuit; and calculates third data according to the column of the first data and the row of the second data.

[0034] Further, the position information of the first data includes: the column coordinates of the first data in the data matrix.

[0035] In a second aspect, an embodiment of the present disclosure provides a matrix calculation method, including:

[0036] Reading and caching the first data in the first matrix and the position information of the first data, where the first matrix is a compressed matrix of the data matrix;

[0037] Reading and caching the second data in the second matrix according to the position information of the first data;

[0038] Calculating third data according to the first data and the second data.

[0039] In a third aspect, an embodiment of the present disclosure provides a processing core, including the matrix calculation circuit, a decoding unit, and a storage device according to any one of the first aspect.

[0040] In a fourth aspect, an embodiment of the present disclosure further provides a chip, where the chip includes at least one processing core in the above third aspect.

[0041] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: a memory for storing computer-readable instructions; and one or more processors for running the computer-readable instructions, so that when the processor runs, it implements any one of the matrix calculation methods in the foregoing first aspect.

[0042] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute any one of the matrix calculation methods in the foregoing first aspect.

[0043] Seventh aspect, an embodiment of the present disclosure provides a computer program product, including computer instructions, when the computer instructions are executed by a computing device, the computing device can execute any of the matrix calculation methods in the foregoing first aspect.

[0044] Eighth aspect, an embodiment of the present disclosure provides a computing device, including one or more chips described in the foregoing fourth aspect.

[0045] An embodiment of the present disclosure discloses a matrix calculation circuit, method, electronic device, and computer-readable storage medium. The matrix calculation circuit includes: a first data reading circuit for reading and caching first data in a first matrix and position information of the first data, where the first matrix is a compressed matrix of a data matrix; a second data reading circuit for reading and caching second data in a second matrix according to the position information of the first data; and a calculation circuit for calculating third data according to the first data and the second data. The above matrix calculation circuit controls the reading of the second data through the position information of the read first data, and solves the technical problems in the prior art that only single data calculation can be performed during matrix calculation and the calculation of the fetch address is complex.

[0046] The above description is only an overview of the technical solutions of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the drawings as follows. Description of the Drawings

[0047] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages, and aspects of each embodiment of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0048] Figures 1a - 1c Schematic diagram of the prior art of the present disclosure;

[0049] Figure 2 Schematic diagram of the structure of the matrix calculation circuit provided by the embodiment of the present disclosure;

[0050] Figure 3 Schematic diagram of the structure of the first data reading circuit provided by the embodiment of the present disclosure;

[0051] Figure 4 Schematic diagram of an example of reordering of the first data reading circuit provided by the embodiment of the present disclosure;

[0052] Figure 5 Schematic diagram of the structure of the second data reading circuit provided by the embodiment of the present disclosure;

[0053] Figures 6a - 6e Schematic diagram of an application example of an embodiment of the present disclosure;

[0054] Figure 7 Flowchart of the matrix calculation method provided by an embodiment of the present disclosure. Detailed implementation manners

[0055] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0056] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0057] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0058] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.

[0059] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0060] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0061] Figure 2 Schematic diagram of the matrix calculation circuit provided by an embodiment of the present disclosure. The matrix calculation circuit (EU) 200 provided in this embodiment includes:

[0062] The first data reading circuit (LD_M1) 201 is configured to read and cache the first data in the first matrix and the position information of the first data, where the first matrix is a compressed matrix of the data matrix;

[0063] The second data reading circuit (LD_M2) 202 is configured to read and cache the second data in the second matrix according to the position information of the first data;

[0064] The calculation circuit 203 is configured to calculate the third data according to the first data and the second data.

[0065] Exemplarily, the first data reading circuit reads and caches a plurality of first data in the first matrix according to the reading address of the first data, and the reading address of the first data is generated according to the storage start address of the first matrix; the second data reading circuit obtains the reading address of the second data according to the position information of the first data, and reads and caches a plurality of second data in the second matrix according to the reading address of the second data, and the reading address of the second data is generated according to the storage start address of the second matrix. Wherein, the storage start address of the first matrix and the storage start address of the second matrix are obtained by the instruction decoding circuit ID (Instruction Decoder), and the instruction decoding circuit is configured to decode the matrix calculation instruction to obtain parameters such as the storage start address of the first matrix, the storage start address of the second matrix, and the sizes of the first matrix and the second matrix.

[0066] Exemplarily, the matrix calculation instruction includes an instruction type, the storage start address of the first matrix, the storage start address of the second matrix, the size of the data matrix, the size of the first matrix, and the size of the second matrix parameters. In one embodiment, the instruction type is a matrix multiplication instruction, the first matrix is a compressed matrix of the data matrix in neural network convolution calculation, and the second matrix is a parameter matrix in neural network convolution calculation; wherein, the data matrix and / or the second matrix is a sparse matrix, and there are a large number of elements in the sparse matrix with a value of 0. It can be understood that the storage start address of the matrix and the size parameters of the matrix (such as the number of rows and columns of the matrix) in the matrix calculation instruction can be represented in the form of register addresses, and the instruction decoding circuit obtains the corresponding data from the corresponding register addresses.

[0067] In an embodiment of the present disclosure, the first data reading circuit 201 receives the starting address of the first matrix decoded by the instruction decoding circuit, generates a reading address for the first data according to the starting address, and reads out a plurality of first data in the first matrix at one time according to the reading address of the first data. Exemplarily, the maximum number of columns read at one time is preset, for example, k columns, where the number of columns refers to the number of columns in the data matrix. Then, the first data reading circuit generates a reading address for the first data according to the starting address of the first matrix and K, and reads out and caches a plurality of first data corresponding to the number of columns and the position information of the plurality of first data in the first matrix at one time. Wherein, the first data is non-zero data in the data matrix, and the first data read by the first data reading circuit is non-zero data in the K columns of data in the data matrix.

[0068] In an embodiment of the present disclosure, the second data reading circuit 202 receives the starting address of the second matrix decoded by the instruction decoding circuit, generates a reading address for the second data according to the starting address and the position information of the plurality of first data, and reads out a plurality of second data in the second matrix at one time according to the reading address of the second data. Exemplarily, the position information of the first data is the column information of the first data in the data matrix, and the column information corresponds to the rows in the second matrix. The row address of the second data to be read is obtained through the starting address of the second matrix and the column information, so that one or more rows of second data can be read out and cached in the second data reading circuit at one time.

[0069] In an embodiment of the present disclosure, the calculation circuit receives the first data transmitted from the first data reading circuit and the second data transmitted from the second data reading circuit, and calculates to obtain third data, where the third data is one or more.

[0070] As Figure 3 shown, in order to implement the functions of the above first data reading circuit, optionally, the first data reading circuit further includes:

[0071] A first data caching circuit 301, a first data sorting circuit 302, and a first control circuit 303;

[0072] Wherein, the first control circuit 303 is used to generate a first data reading address according to the starting address of the first matrix;

[0073] The first data caching circuit 301 is used to cache the first data read according to the first data reading address and the position information of the first data;

[0074] The first data sorting circuit 302 is configured to re-sort the first data position information and the first data in a one-to-one correspondence manner according to the first data position information in the first data caching circuit, where the re-sorting result is that the data in the same row of the data matrix remains in the same row, and the data in the same column of the data matrix remains in the same column.

[0075] Optionally, the first control circuit 303 receives the starting address of the first matrix decoded by the instruction decoding circuit, a preset parameter K, and the size parameter of the first matrix, such as the non-zero data in the N columns of the data matrix included in the first matrix. Optionally, the first control circuit includes a first reading control circuit CL1 and a first address generation circuit AG1. The first reading control circuit CL1 receives the starting address of the first matrix decoded by the instruction decoding circuit, the preset parameter K, and the size parameter of the data matrix, etc., and controls AG1 to generate a first data reading address Addr1, so that the first data reading circuit can read the first data corresponding to the k-th column of the data matrix in the first matrix according to the Addr1 at one time.

[0076] Optionally, the first data caching circuit 301 further includes a first memory or a first storage area DB11 for caching the first data, and a second memory or a second storage area DB10 for caching the position information of the first data. After reading the first data and the position information of the first data from the first matrix, the first data is cached in DB11, and the position information of the first data is cached in DB10.

[0077] Optionally, the first data sorting circuit 302 further includes a re-sorted position information caching circuit IRDB and a re-sorted first data caching circuit DRDB. The IRDB is used to cache the position information of the re-sorted first data, and the DRDB is used to cache the re-sorted first data.

[0078] Optionally, the position information of the first data includes the row coordinate and column coordinate of the first data in the data matrix. The row coordinate is represented by the X coordinate, and the column coordinate is represented by the Y coordinate. Exemplarily, for the reordering of the data, the reordering can be performed in the order of columns first and then rows, that is, first in ascending order of the Y coordinate, and then in ascending order of the X coordinate to ensure that the first data in the same row of the data matrix is still in the same row, and the first data not in the same row is still not in the same row. It also ensures that the first data in the same column of the data matrix is still in the same column, and the first data not in the same column is still not in the same column. Since the first matrix is the compression matrix of the data matrix, some rows in the data matrix lack the first data of this column, while other rows have the data of this column. Then, when reordering, 0 will be filled in the position of this row and this column. The reordered Y coordinates are cached in the IRDB, and the reordered first data is cached in the DRDB.

[0079] Figure 4 For the schematic diagram of the reordering example, as Figure 4 shown, the data matrix M1_O is a sparse matrix, and the first matrix is the compression matrix M1 of the data matrix. M1 includes the first data Data in the data matrix and the position information (X, Y) of the first data in the data matrix. When the first data reading circuit reads 2 columns of data in M1, the 2 columns of data are reordered according to the position information of the 2 columns of data. For example, it can be reordered in the order of first arranging in ascending order of the Y coordinate and then in ascending order of the X coordinate, and the position information with the same X coordinate is in the same row, and the position information with different X coordinates is in different rows, and the position information with the same Y coordinate is in the same column, and the position information with different Y coordinates is in different columns. The first data is stored at the position corresponding to the sorting of the position information. At the same time, the column information in the position information corresponding to the 2 columns of data is saved, that is, only the Y coordinate is saved and stored in ascending order of the Y coordinate.

[0080] As Figure 4 shown, after the reordering, the first data 1 is in the first row, and the first data 2 is in the second row after reordering, and 0 is filled in other positions; and the Y-axis coordinates in the position information are saved in ascending order.

[0081] After reordering, the first data reading circuit outputs the position information DO0 and the first data DO1. Wherein DO1 is part or all of the first data in the first data, and the position information DO0 is the position information of all the first data in the DRDB. The position information is transmitted to the second data reading circuit so that the second data reading circuit reads one or more second data corresponding to the position information.

[0082] As Figure 5As shown, in order to implement the functions of the above-mentioned second data reading circuit, optionally, the second data reading circuit further includes:

[0083] A second data cache circuit 501 and a second control circuit 502;

[0084] Among them, the second control circuit 502 is used to generate a second data reading address according to the starting address of the second matrix and the position information of the first data;

[0085] The second data cache circuit 501 is used to cache the second data read out according to the second data reading address.

[0086] Optionally, the second control circuit 502 receives the starting address of the second matrix decoded by the instruction decoding circuit and the column information in the position information of the first data to generate a second data reading address. Optionally, the second control circuit includes a second reading control circuit CL2 and a second address generation circuit AG2. The second reading control circuit CL2 receives the starting address of the second matrix decoded by the instruction decoding circuit and the position information of the first data, and controls AG2 to generate a second data reading address Addr2, so that the second data reading circuit can read the second data corresponding to the first data in the second matrix according to the Addr2 at one time. Exemplarily, the position information of the first data is the column coordinate of the first data. According to the starting address of the second matrix and the column coordinate, the second address generation circuit AG2 uses the starting address of the second matrix as the base address, and uses the column coordinate as the row offset value of the second data to obtain the row address of the second data, so that one or more rows of second data corresponding to the first data can be read. Optionally, the first data is the first data corresponding to K columns in the data matrix, and the second data is the second data corresponding to the K rows in the second matrix corresponding to the K columns of the first data in the matrix calculation.

[0087] Optionally, the second data cache circuit 501 includes a second data memory or a second data storage area, the size of which is the size of the second data of K rows, and the read second data is cached row by row in the second data cache circuit according to its position in the second matrix.

[0088] As Figure 2 shown, the computing circuit 203 includes:

[0089] A computing unit array PUA, and the computing unit array includes a plurality of computing units PU 1,1 , PU 1,2 , …… PU M,N ;

[0090] The row computing units in the computing unit array simultaneously receive a row of the second data in the second data;

[0091] The column computing units in the computing unit array simultaneously receive a column of the first data in the first data.

[0092] Optionally, the computing circuit 203 receives a column of the first data output by the first data sorting circuit; receives a row of the second data output by the second data caching circuit; and calculates a third data based on the column of the first data and the row of the second data. Specifically, one of the first data in the column of the first data output by the first data sorting circuit is output to a column of computing units in the computing circuit. For example, if the column of the first data includes two first data, the first data in the 0th row of the column of the first data is output to the computing unit in the 0th row of this column of computing units, and the first data in the 1st row of the column of the first data is output to the computing unit in the 1st row of this column of computing units. The above operation is performed for each column of computing units participating in the calculation, and the first data is used as the first input data of the computing unit; the row of the second data output by the second data caching circuit is output to a row of computing units in the computing unit array. Specifically, the second data caching circuit outputs a row of the second data corresponding to the column of the first data output by the first data sorting circuit. For example, if the column of the first data includes 2 first data, the second data output by the second data caching circuit is a row of the second data, and the row of the second data includes 2 second data. The 2 second data in this row are respectively input into the corresponding row of computing units, that is, the 1st second data is output to the 1st computing unit in the row of computing units, and the 1st second data is output to the 2nd computing unit in the row of computing units; thus, each computing unit participating in the calculation will obtain two data inputs, a first data and a second data. The computing unit calculates the calculation result of the first data and the second data through the calculation type specified by the type of the calculation instruction to obtain the third data. Multiple computing units obtain the third data and output it. The above calculation process is looped, and each computing unit accumulates its calculation result until all the first data and the second data are read out to obtain an output matrix, where the value of each element in the output matrix is the accumulated result of the computing units participating in the calculation.

[0093] Figures 6a - 6e It is an example of the calculation process of the matrix calculation circuit in the above embodiment. As Figure 6a shown, it is the matrix multiplication calculation that the matrix calculation circuit needs to perform. M1_O is the data matrix, M2 is the second matrix, and M is the third matrix M obtained by multiplying M1_O and M2 matrices.

[0094] Among them, M1_O is stored in the form of a compressed matrix. As Figure 6bAs shown, compress M1_0 to generate the first matrix M1 and save it. Let K = 4, that is, during the calculation process, each time 4 columns of the first data in the data matrix M1_O are read. For this example, all the data in M1 is read and cached at one time. Then as Figure 6b shown, the first data reading circuit of the matrix calculation circuit reads the first data in the entire first matrix M1 into the data cache circuit at one time, and after being re-ordered by the first data sorting circuit, the storage order in the IRDB and DRDB as shown in Figure 6b is obtained.

[0095] As Figure 6c shown is the overall schematic diagram of matrix calculation using the matrix calculation circuit. Read 4 columns of the first data of M1 in units of K = 4 columns, that is, the columns with column numbers 0 - 3 in the data matrix. Since in this example, the total number of columns of the data matrix M1_O is 4, the entire M1 will be read and cached into the first data reading circuit LD_M1 at one time; after reading, re-ordering is performed, the position information is stored in the IRDB of LD_M1, and the first data is stored in the DRDB of LD_M1. The IRDB transmits the column coordinates 0 and 3 to LD_M2, and LD_M2 reads the 0th row and 3rd row of the second data from M2 according to the column coordinates 0 and 3 and caches them, as Figure 6c shown, the second data in the 0th row includes 1 and 2, and the second data in the 3rd row includes 7 and 8. These two rows of second data are cached in the DB2 of LD_M2.

[0096] As Figure 6d shown is the schematic diagram of the first calculation. The calculation circuit obtains the first column of the first data from the DRDB of LD_M1. The first column of the first data includes 1 in the 0th row and 0 in the 1st row. Among them, 1 in the 0th row is input into the 0th row calculation units PU 0,0 and PU 0,1 in; 0 in the 1st row is input into the 1st row calculation units PU 1,0 and PU 1,1 in; the second data cache circuit of LD_M2 outputs the second data in the 0th row cached in LD_M2. The second data in the 0th row is input into the 0th row calculation units PU 0,0 and PU 0,1 in and the 1st row calculation units PU 1,0 and PU 1,1 in. The second data in the 0th row includes 1 and 2. Among them, the second data 1 is input into the calculation unit PU 0,0 and the calculation unit PU 1,0 in, and the second data 2 is input into the calculation unit PU 0,1 and PU 1,1 in. After that, each calculation unit independently performs multiply-accumulate calculations and respectively obtains PU 0,0Calculation result 1, PU 0,1 Calculation result 2, PU 1,0 Calculation result 0 and PU 1,1 Calculation result 0; since the first data and the second data have not been calculated completely, the third data obtained is the intermediate data M_temp.

[0097] Such as Figure 6e Shown is a schematic diagram of the second calculation. The calculation circuit obtains the first data in the second column from the DRDB of LD_M1, where the first data in the second column includes 0 in the 0th row and 2 in the 1st row, and 0 in the 0th row is input into the 0th row calculation unit PU 0,0 and PU 0,1 ; 2 in the 1st row is input into the 1st row calculation unit PU 1,0 and PU 1,1 ; The second data cache circuit of LD_M2 outputs the second data in the 1st row cached in LD_M2, where the second data in the 1st row includes 7 and 8, and the second data 7 is input into the calculation unit PU 0,0 and PU 1,0 ; The second data 8 is input into the calculation unit PU 0,1 and PU 1,1 ; Then each calculation unit independently performs multiply-accumulate calculations, and respectively obtains the calculation result 1 of PU 0,0 the calculation result 2 of PU 0,1 the calculation result 14 of PU 1,0 and the calculation result 16 of PU 1,1 ; Since the first data and the second data have been calculated, the third data obtained is the value of the element in the output matrix M.

[0098] From the calculation process of the above example, it can be seen that using the matrix calculation circuit in the present disclosure to perform matrix multiplication only requires two calculations to complete the multiplication of a 2*4 matrix and a 4*2 matrix, greatly improving the calculation speed and saving the calculation time.

[0099] Through the above technical solution of the present disclosure, directly calculate the compressed sparse matrix, effectively saving storage space and data bandwidth; using the calculation unit array, all calculation units perform data processing synchronously, greatly improving the data utilization rate, and multiple calculation units can share the same data; directly calculate the compressed sparse matrix, skipping the calculation of some 0 elements, thereby improving the operation speed and the effective computing power of the chip.

[0100] Figure 7 This is the flowchart of the matrix calculation method provided by the embodiment of the present disclosure. Such as Figure 7 shown, the method includes the following steps:

[0101] Step S701, read and cache the first data in the first matrix and the position information of the first data, where the first matrix is a compressed matrix of a data matrix;

[0102] Step S702, read and cache the second data in the second matrix according to the position information of the first data;

[0103] Step S703, calculate the third data according to the first data and the second data;

[0104] Further, the reading and caching the first data in the first matrix and the position information of the first data includes:

[0105] Generate a first data reading address according to the starting address of the first matrix;

[0106] Cache the first data read out according to the first data reading address and the position information of the first data;

[0107] Re - sort the position information of the first data and the first data in a one - to - one corresponding manner according to the position information of the first data. Among them, the result of the re - sorting is that the data in the same row of the data matrix is still in the same row, and the data in the same column of the data matrix is still in the same column.

[0108] Further, the method further includes:

[0109] Send the position information of the first data.

[0110] Further, the reading and caching the second data in the second matrix according to the position information of the first data includes:

[0111] Generate a second data reading address according to the starting address of the second matrix and the position information of the first data;

[0112] Cache the second data read out according to the second data reading address.

[0113] Further, the generating a second data reading address according to the starting address of the second matrix and the position information of the first data includes:

[0114] Generate the second data reading address according to the starting address of the second matrix and the column information in the position information of the first data.

[0115] Further, the calculating the third data according to the first data and the second data includes:

[0116] Receive a sorted column of first data; receive a row of second data corresponding to the column of first data in the second data; calculate third data based on the sorted column of first data and the row of second data.

[0117] Further, the position information of the first data includes: the column coordinate of the first data in the data matrix.

[0118] In the foregoing, although the steps in the above method embodiments are described in the above order, those skilled in the art should understand that the steps in the embodiments of the present disclosure do not necessarily need to be executed in the above order, and they can also be executed in reverse order, in parallel, in a cross-over manner, or other orders. Moreover, based on the above steps, those skilled in the art can also add other steps, and these obvious variations or equivalent replacement methods should also be included in the protection scope of the present disclosure, and will not be elaborated herein.

[0119] The embodiments of the present disclosure further provide a processing core, which includes at least one matrix calculation circuit, a decoding unit, and a storage device in the above embodiments.

[0120] The embodiments of the present disclosure further provide a chip, which includes at least one processing core in the above embodiments.

[0121] The embodiments of the present disclosure provide an electronic device, including: a memory for storing computer-readable instructions; and one or more processors for running the computer-readable instructions, so that when the processor runs, it implements any of the matrix calculation methods in the embodiments.

[0122] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium, which is characterized in that the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute any of the matrix calculation methods in the foregoing embodiments.

[0123] The embodiments of the present disclosure further provide a computer program product, which is characterized in that it includes computer instructions, and when the computer instructions are executed by a computing device, the computing device can execute any of the matrix calculation methods in the foregoing embodiments.

[0124] The embodiments of the present disclosure further provide a computing device, which is characterized in that it includes any of the chips in the embodiments.

[0125] The flowcharts and block diagrams in the accompanying drawings of the present disclosure illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a task, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0126] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0127] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0128] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

Claims

1. A matrix calculation circuit, characterized in that, Comprising: A first data reading circuit, configured to read and cache first data in a first matrix and position information of the first data, wherein the first matrix is a compressed matrix of a data matrix; A second data reading circuit, configured to read and cache second data in a second matrix according to the position information of the first data; A calculation circuit, configured to calculate third data according to the first data and the second data; The first data reading circuit further includes: A first data caching circuit, a first data sorting circuit, and a first control circuit; Wherein, the first control circuit is configured to generate a first data reading address according to a starting address of the first matrix; The first data caching circuit is configured to cache the first data read according to the first data reading address and the position information of the first data; The first data sorting circuit is configured to re-sort the position information of the first data and the first data in a one-to-one position correspondence manner according to the position information of the first data in the first data caching circuit, wherein, the re-sorting result is that the data in the same row of the data matrix is still in the same row, and the data in the same column of the data matrix is still in the same column.

2. The matrix calculation circuit according to claim 1, wherein, The first data sorting circuit is further configured to: Send the position information of the first data to the second data reading circuit.

3. The matrix calculation circuit according to any one of claims 1-2, wherein, The second data reading circuit further includes: A second data caching circuit and a second control circuit; Wherein, the second control circuit is configured to generate a second data reading address according to a starting address of the second matrix and the position information of the first data; The second data caching circuit is configured to cache the second data read according to the second data reading address.

4. The matrix calculation circuit according to claim 3, wherein the second control circuit is configured to generate a second data reading address according to a starting address of the second matrix and the position information of the first data, including: The second control circuit is configured to generate the second data reading address according to the starting address of the second matrix and column information in the position information of the first data.

5. The matrix calculation circuit according to claim 1, wherein the calculation circuit includes: An array of calculation units, wherein the array of calculation units includes a plurality of calculation units; Row calculation units in the array of calculation units simultaneously receive a row of second data in the second data; Column calculation units in the array of calculation units simultaneously receive a column of first data in the first data.

6. The matrix calculation circuit according to claim 3, wherein the calculation circuit is configured to calculate third data according to the first data and the second data, including: The calculation circuit receives a re-sorted column of first data output by the first data sorting circuit; Receives a row of second data output by the second data caching circuit; Calculates third data according to the re-sorted column of first data and the row of second data.

7. The matrix calculation circuit according to any one of claims 1-2, 4-6, wherein the position information of the first data includes: The column coordinate of the first data in the data matrix.

8. A matrix calculation method, characterized in that, Comprising: Read and cache the first data in the first matrix and the position information of the first data, where the first matrix is a compressed matrix of a data matrix; Read and cache the second data in the second matrix according to the position information of the first data; Calculate the third data based on the first data and the second data; The reading and caching of the first data in the first matrix and the position information of the first data include: Generate a first data reading address according to the starting address of the first matrix; Cache the first data read out according to the first data reading address and the position information of the first data; Reorder the first data position information and the first data in a one-to-one correspondence manner according to the first data position information in the first data cache circuit. Among them, the reordering result is that the data in the same row of the data matrix is still in the same row, and the data in the same column of the data matrix is still in the same column.

9. A processing core, comprising the matrix calculation circuit according to any one of claims 1-7.

Citation Information

Patent Citations

  • Calculation engine and electronic equipment

    CN106126481A

  • Computing device and related product

    CN111353591A