A Fast Calculation Method, Device, Equipment and Storage Medium for Matrix Multiplication
By compressing and transposing the matrix and using the finite domain operation method, the optimization problem of matrix multiplication in the finite domain is solved, and the rapid calculation and efficient performance of matrix multiplication are achieved.
Patent Information
- Application Number
- CN202510116827.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the prior art, the optimization algorithm of matrix multiplication in the finite domain has not been effectively solved, resulting in large resource consumption and large delays, especially in cache memory.
By compressing and transposing the matrix, using the actual bit width of the processor to perform matrix multiplication operations, and using the exclusive or sum of finite domains to achieve fast calculation of matrix multiplication.
It improves the cache hit rate, reduces the waste of hardware resources, improves computing performance, and avoids the need for restore processing of compressed matrices.
Smart Images

Figure CN119557552B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of electronic digital data processing, and particularly to a method, apparatus, device, and storage medium for fast matrix multiplication calculation. Background Art
[0002] Matrix multiplication is a core technology for data processing and is widely used in the fields of communication, image processing, and scientific computing. With the development of technologies such as big data and artificial intelligence, the performance and power consumption of matrix operations have become very important. Large matrix multiplications consume a large amount of resources and have a large latency. To calculate matrix multiplications, various optimization schemes have been proposed. For matrices with different characteristics, the proposed optimization schemes are also different. Finite field matrices have special elements, where each element in such a matrix is an element in a finite field, and the operations of the elements also follow the operation rules of the finite field. Matrix operations in finite fields are more common in communication, coding, and encryption systems.
[0003] In the prior art, whether the matrix is stored in memory in a row-major storage format or a column-major storage format, when the cache fetches a row of matrix A or a column of matrix B from the main memory, due to discontinuous addresses, there will be a large number of cache misses on one side. One method is to perform matrix block operations, but there is still a problem of high complexity. In the prior art, no optimized algorithm for matrix multiplication in finite fields has been found. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and storage medium for fast matrix multiplication calculation to at least solve the above technical problems existing in the prior art.
[0005] According to a first aspect of the present disclosure, a method for fast matrix multiplication calculation is provided, the method including:
[0006] Based on different storage positions, respectively determine a first matrix and a third matrix;
[0007] Obtain the actual bit width of the processor, and compress the first matrix and the third matrix based on the actual bit width to obtain a fourth matrix and a fifth matrix;
[0008] Perform finite field multiplication and addition operations on the fourth matrix and the fifth matrix using matrix multiplication to obtain a seventh matrix.
[0009] In an implementable embodiment, the step of respectively determining the first matrix and the third matrix based on different storage positions includes:
[0010] Based on different storage positions, respectively determine the first matrix and a second matrix;
[0011] Read the storage formats of the first matrix and the second matrix respectively;
[0012] Determine the dimension information of the first matrix and the second matrix based on the storage formats, and judge whether the first matrix and the second matrix conform to the matrix multiplication operation rules based on the dimension information;
[0013] If they do not conform to the matrix multiplication operation rules, generate an error message;
[0014] If they conform to the matrix multiplication operation rules, then transpose the storage format of the second matrix based on the matrix multiplication rules to obtain a third matrix.
[0015] In an implementable manner, the obtaining of the actual bit width of the processor includes:
[0016] Measure the actual number of bytes of a single data based on a first keyword, and calculate the actual bit width of the processor based on the actual number of bytes of the single data.
[0017] In an implementable manner, the method further includes:
[0018] Determine a first data volume based on the memory sizes occupied by the first matrix and the second matrix;
[0019] Calculate the memory required for the fourth matrix and the fifth matrix based on the actual bit width and the first data volume, determine a second data volume based on the required memory, and allocate memory spaces for the fourth matrix and the fifth matrix based on the second data volume.
[0020] In an implementable manner, the compressing the first matrix and the third matrix based on the actual bit width includes:
[0021] Divide the column indexes of the first matrix into a first bit index and a first other column index;
[0022] Read the first bit index therein based on the column indexes of the first matrix, and calculate the bit index of the compressed fourth matrix based on the first bit index and the actual bit width;
[0023] Read the first other column index therein based on the column indexes of the first matrix, and calculate the other column index of the compressed fourth matrix based on the first other column index and the actual bit width;
[0024] Construct the fourth matrix based on the bit index and the other column index of the fourth matrix;
[0025] Divide the indexes of the third matrix into a third bit index and a third other index;
[0026] Read the third bit index therein based on the column index of the third matrix, and calculate the bit index of the compressed fifth matrix based on the third bit index and the actual bit width;
[0027] Read the third other column index therein based on the column index of the third matrix, and calculate the other column index of the compressed fifth matrix based on the third other column index and the actual bit width;
[0028] Construct the fifth matrix based on the bit index of the fifth matrix and the other column index of the fifth matrix.
[0029] In an implementable manner, the fourth matrix and the fifth matrix perform a finite field multiply-add operation based on matrix multiplication to obtain a seventh matrix, including:
[0030] Perform a finite field multiply-add operation on the fourth matrix and the fifth matrix using matrix multiplication to obtain a sixth matrix;
[0031] Based on finite field addition, add each bit of data in the sixth matrix to obtain the seventh matrix.
[0032] In an implementable manner, the fourth matrix and the fifth matrix perform a finite field multiply-add operation based on matrix multiplication to obtain a sixth matrix, including:
[0033] Select the corresponding rows in the fourth matrix and the fifth matrix, and perform a bitwise AND operation on the selected corresponding elements in the corresponding rows;
[0034] Perform an exclusive OR operation on the elements after the bitwise AND operation to obtain the sixth matrix.
[0035] According to the second aspect of the present disclosure, there is provided a matrix multiplication fast calculation device, the device includes:
[0036] A data input / output module, configured to respectively determine a first matrix and a third matrix based on different storage locations;
[0037] A matrix compression module, configured to obtain the actual bit width of the processor, and compress the first matrix and the third matrix based on the actual bit width to obtain a fourth matrix and a fifth matrix;
[0038] A matrix operation module, configured to perform a finite field multiply-add operation on the fourth matrix and the fifth matrix using matrix multiplication to obtain a seventh matrix.
[0039] In an implementable manner, the device further includes:
[0040] A parameter configuration module, which is used to respectively determine a first matrix and a second matrix based on different storage locations; respectively read the storage formats of the first matrix and the second matrix; determine the dimension information of the first matrix and the second matrix based on the storage formats, and judge whether the first matrix and the second matrix conform to the matrix multiplication operation rules based on the dimension information; if they do not conform to the matrix multiplication operation rules, generate an error message.
[0041] A matrix transpose module, which is used to transpose the storage format of the second matrix based on the matrix multiplication rule to obtain a third matrix.
[0042] The data input / output module is further used to determine a first data volume based on the memory sizes occupied by the first matrix and the second matrix; calculate the memory required for the fourth matrix and the fifth matrix based on the actual bit width and the first data volume, determine a second data volume based on the required memory, and allocate memory spaces for the fourth matrix and the fifth matrix based on the second data volume. In an implementable embodiment, the matrix compression unit further includes:
[0043] A processor bit width calculation module, which is used to measure the actual number of bytes of a single data based on a first keyword, and calculate the actual bit width of the processor based on the actual number of bytes of the single data.
[0044] In an implementable embodiment, the matrix compression module is further used to divide the column indexes of the first matrix into a first bit index and a first other column index; read the first bit index therein based on the column indexes of the first matrix, and calculate the bit index of the compressed fourth matrix based on the first bit index and the actual bit width; read the first other column index therein based on the column indexes of the first matrix, and calculate the other column index of the compressed fourth matrix based on the first other column index and the actual bit width; construct the fourth matrix based on the bit index and the other column index of the fourth matrix; divide the indexes of the third matrix into a third bit index and a third other index; read the third bit index therein based on the column indexes of the third matrix, and calculate the bit index of the compressed fifth matrix based on the third bit index and the actual bit width; read the third other column index therein based on the column indexes of the third matrix, and calculate the other column index of the compressed fifth matrix based on the third other column index and the actual bit width; construct the fifth matrix based on the bit index and the other column index of the fifth matrix.
[0045] In an implementable embodiment, the matrix operation module is further used to perform a finite field multiplication and addition operation on the fourth matrix and the fifth matrix by using matrix multiplication to obtain a sixth matrix.
[0046] The matrix operation module includes: a matrix accumulation module, which is used to add each bit of data in the sixth matrix based on the addition in the finite field to obtain the seventh matrix.
[0047] In an implementable embodiment, the matrix operation module further includes: a matrix multiplication module, which is used to select corresponding rows in the fourth matrix and the fifth matrix, and perform a bitwise AND operation on the selected corresponding elements in the corresponding rows;
[0048] The matrix accumulation module is further used to perform an exclusive OR operation on the elements after the bitwise AND operation to obtain the sixth matrix.
[0049] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0050] At least one processor; and
[0051] A memory communicatively connected to the at least one processor; wherein,
[0052] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the present disclosure.
[0053] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the method described in the present disclosure.
[0054] A method, device, equipment and storage medium for fast matrix multiplication according to the present disclosure, before performing matrix multiplication, by transposing the matrix stored in row-major order, adjusts the storage order of matrix multiplication, improves the cache hit rate, and performs matrix compression, and performs exclusive OR operation and AND operation on the compressed matrix to obtain the calculation results of the two matrices before transposition. A scheme of compressing first and then operating is proposed and optimized, which greatly improves the utilization rate of hardware resources, converts the addition, subtraction, multiplication and division operations in the rational number field into exclusive OR and AND operations in the finite field, greatly improves the operation performance, and at the same time, adjusts the calculation order and no longer needs to restore the compressed matrix.
[0055] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0056] By reading the following detailed description with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, wherein:
[0057] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0058] Figure 1 The implementation process of a fast matrix multiplication calculation method according to an embodiment of the present disclosure is shown schematically Figure 1 ;
[0059] Figure 2 The implementation process of a fast matrix multiplication calculation method according to an embodiment of the present disclosure is shown schematically Figure 2 ;
[0060] Figure 3 The implementation process of a fast matrix multiplication calculation method according to an embodiment of the present disclosure is shown schematically Figure 3 ;
[0061] Figure 4 Shows a schematic diagram of the overall implementation scheme of matrix multiplication in the embodiments of the present disclosure;
[0062] Figure 5 Shows a schematic diagram of the specific implementation scheme of matrix compression in the embodiments of the present disclosure;
[0063] Figure 6 Shows a schematic diagram of the specific implementation scheme of grouped calculation of matrix elements in the embodiments of the present disclosure;
[0064] Figure 7 Shows a schematic diagram of a fast matrix multiplication calculation device according to an embodiment of the present disclosure Figure 1 ;
[0065] Figure 8 Shows a schematic diagram of a fast matrix multiplication calculation device according to an embodiment of the present disclosure Figure 2 ;
[0066] Figure 9 Shows a schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0067] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.
[0068] In the prior art, the finite field GF(2) has only two elements, 0 and 1, and follows the following operation rules: a + b = a - b = a ^ b, and both the addition operation and the subtraction operation can be converted into the exclusive OR operation; a × b = a & b, and the multiplication operation can be converted into the AND operation.
[0069] Suppose A is an m × n matrix, and B is an n × p matrix. When multiplying the two matrices, the resulting matrix C is an m × p matrix, and follows the following rules:
[0070]
[0071] That is, the element at index (i, j) in the C matrix is equal to the sum of the products of all elements in the i-th row of the A matrix and all elements in the j-th column of the B matrix. Generally, when calculating matrix multiplication, it is necessary to traverse the rows of the A matrix, the columns of the B matrix, and the corresponding elements of the two rows and columns. Therefore, due to the closure of the finite field operation, matrix multiplication can be processed in parallel.
[0072] Figure 1 shows the implementation process schematic of a fast matrix multiplication calculation method according to an embodiment of the present disclosure Figure 1 ; as Figure 1 shown, a fast matrix multiplication calculation method according to an embodiment of the present disclosure includes the following steps:
[0073] Step 101, respectively determine a first matrix and a third matrix based on different storage locations.
[0074] In the embodiment of the present disclosure, a first matrix and a second matrix are respectively determined based on different storage locations; wherein, the matrix generally adopts a row-major storage format, and the storage formats of the first matrix and the second matrix are respectively read; based on the storage format, the dimension information of the first matrix and the second matrix is determined, and based on the dimension information, it is determined whether the first matrix and the second matrix conform to the matrix multiplication operation rule; if they do not conform to the matrix multiplication operation rule, an error message is generated, where the error message is used to prompt the system to report an error; if they conform to the matrix multiplication operation rule, since matrix calculation requires reading an entire row of the first matrix and an entire column of the second matrix respectively, in order to improve the cache hit rate, based on the matrix multiplication rule, the storage format of the second matrix is transposed to obtain a third matrix.
[0075] In the embodiment of the present disclosure, the actual byte number of a single data is measured based on a first keyword, where the first keyword can be: the keyword of the data type length function (sizeof), and the actual bit width of the processor is calculated based on the actual byte number of the single data, specifically: multiplying the measured actual byte number of the single data by 8 is the actual bit width.
[0076] In the embodiments of the present disclosure, obtain the memory sizes occupied by the storage of the first matrix and the second matrix, denoted as the first data volume; calculate the memory required for the fourth matrix and the fifth matrix based on the actual bit width and the first data volume, denoted as the second data volume, and allocate memory spaces for the fourth matrix and the fifth matrix based on the second data volume. Among them, since the fourth matrix and the fifth matrix are the matrices obtained by compressing the first matrix and the third matrix, and the matrix compression is based on the actual bit width of the processor, therefore, according to the data volumes of the first matrix and the third matrix before compression, that is, the size of the first data volume, the data volumes of the compressed fourth matrix and fifth matrix can be calculated using the actual bit width, which is the second data volume. Allocate corresponding memory spaces according to the size of the second data volume.
[0077] Step 102: Obtain the actual bit width of the processor, and compress the first matrix and the third matrix based on the actual bit width to obtain a fourth matrix and a fifth matrix.
[0078] In the embodiments of the present disclosure, determine the corresponding relationship between the matrices before and after compression, and implement matrix compression according to the corresponding relationship between the matrix elements before and after compression. Specifically: divide the column index of the first matrix elements into a first bit index and a first other column index; read the first bit index therein based on the column index of the first matrix as the bit index of the compressed matrix elements, and calculate the bit index of the compressed fourth matrix based on the first bit index and the actual bit width, indicating which bit of the corresponding element in the compressed fourth matrix is the first matrix element before compression; read the first other column index therein based on the column index of the first matrix, that is, the column index remaining after removing the bit index, and calculate the other column index of the compressed fourth matrix based on the first other column index and the actual bit width, representing which elements of the first matrix before compression are included in the elements of the compressed fourth matrix; construct the fourth matrix based on the bit index of the fourth matrix and the other column index of the fourth matrix.
[0079] In the embodiments of the present disclosure, the compression method of the third matrix is the same as that of the first matrix above: divide the index of the third matrix into a third bit index and a third other index; read the third bit index therein based on the column index of the third matrix, and calculate the bit index of the compressed fifth matrix based on the third bit index and the actual bit width; read the third other column index therein based on the column index of the third matrix, and calculate the other column index of the compressed fifth matrix based on the third other column index and the actual bit width; construct the fifth matrix based on the bit index of the fifth matrix and the other column index of the fifth matrix.
[0080] Step 103: Perform finite field multiplication and addition operations on the fourth matrix and the fifth matrix using matrix multiplication to obtain a seventh matrix.
[0081] In the embodiments of the present disclosure, finite field multiplication and addition operations are performed on the fourth matrix and the fifth matrix using matrix multiplication to obtain a sixth matrix. Specifically: corresponding rows in the fourth matrix and the fifth matrix are selected, and the selected corresponding elements in the corresponding rows are subjected to a bitwise AND operation (the AND operation is equivalent to the multiplication operation in the gf(2) finite field); the elements after the bitwise AND operation are subjected to an exclusive OR operation (in the finite field, the addition operation is equivalent to the exclusive OR operation) to obtain the sixth matrix. Based on finite field addition, each bit of data in the sixth matrix is added to obtain the seventh matrix. In the actual implementation process, the initial value of the exclusive OR operation is assigned 0, and each time a result of the AND operation is obtained, it is added to the result of the previous exclusive OR operation using the exclusive OR operation.
[0082] In the embodiments of the present disclosure, the first matrix, the second matrix, the third matrix, the fourth matrix, the fifth matrix, the sixth matrix, and the seventh matrix can be stored in a random access memory (RAM). Among them, matrices with a conversion relationship can be stored in the same RAM. Specifically: the first matrix and the fourth matrix are stored in the same RAM, the second matrix, the third matrix, and the fifth matrix are stored in the same RAM, and the sixth matrix and the seventh matrix are stored in the same RAM.
[0083] Figure 2 Shows the implementation process schematic of a fast matrix multiplication calculation method according to an embodiment of the present disclosure Figure 2 , as Figure 2 shown, a fast matrix multiplication calculation method according to an embodiment of the present disclosure includes the following steps:
[0084] Step 201: Adjust the matrix.
[0085] In the embodiments of the present disclosure, in order to improve the cache hit rate, the matrix storage format needs to be changed first. First, the input matrix A and matrix B are received through the configured interface, and the dimensions of matrix A and matrix B are read. Among them, the number of rows of A is row_A, the number of columns of A is col_A; the number of rows of B is row_B, and the number of columns of B is col_B. To conform to the operation rules of matrix multiplication, it is necessary to ensure that: col_A = row_B.
[0086] In the embodiments of the present disclosure, assuming that the matrix adopts a row-major storage format, calculating an element in the C matrix requires reading an entire row row_A of the A matrix and an entire column col_B of the B matrix. Since it is a row-major storage method, a continuous segment of data in an entire row row_A of matrix A can be loaded into the cache and can all be used for matrix multiplication, that is, the cache hit rate of matrix A is very high. For matrix B, each time a row row_B is read, but only 1 number is used, and the cache hit rate of matrix B is very low. A method to improve the cache hit rate of matrix B is to use it after transposing matrix B. That is, use the transposed matrix of B instead of directly using matrix B when calculating matrix multiplication. Here, the transposed matrix of B is denoted as B_rev. Then, the method of calculating C(i,j) becomes changing the operation method of multiplying the i-th row col_A of matrix A and the j-th column col_B of matrix B to the operation between the i-th row col_A of matrix A and the j-th row row_B of matrix B_rev.
[0087] Step 202, compress the matrix.
[0088] In the embodiments of the present disclosure, since both A and B_rev are row-wise operations, the method of compressing first and then calculating can be used. The principle is as follows: The elements in gf(2) can be represented by 1 bit, and the operations in gf(2) are all completed within this 1 bit without carry phenomenon. However, the processors inside the computer store and operate according to 32 bits or 64 bits. Therefore, the 1-bit operation a^b in gf(2) will be converted into a 32-bit operation in the processor, which will cause a large waste of computing resources in the processor. We assume that the CPU bit width is 32 bits. During the matrix multiplication process, the main operations are performed row by row, which requires operating on entire rows. So the data in a row can be divided into groups of 32 each, and each group of data is compressed into one data in the computer. One bit operation of the processor can complete the operation of 32 data in the matrix, thus making full use of the processor resources and greatly improving the computing performance.
[0089] In the embodiments of the present disclosure, preferably, the storage space with address 1 is regarded as a 32-bit cache space for storing the compressed data. Only the lowest bit in the input data is valid. The lowest bit is intercepted and placed at the corresponding position in the 32-bit cache space. That is, the valid bit of the first data is stored in the 1st bit of the cache space, the valid bit of the second data is placed in the 2nd bit of the cache space, and so on. After the storage space with address 1 is full, it switches to the next address (address 2), regarded as a new 32-bit cache space, to store the next 32 data. Thus, after the matrix input is completed, the compressed data of the matrix will be stored.
[0090] Step 203, calculate in groups.
[0091] In the embodiments of the present disclosure, when calculating each C(i,j), there are a total of 2n numbers, with n multiplication operations and n - 1 addition operations performed. Since the CPU can calculate 32 data at a time, these 2n numbers can be grouped. For example, the n numbers of matrix A are divided into groups of 32 numbers each, and matrix B is grouped in the same way. When calculating matrix multiplication, first calculate the multiplication between the corresponding elements of the groups of matrix A and matrix B. Then, add the 32 elements obtained after multiplying the corresponding elements of these 32 elements with the subsequent corresponding elements according to the corresponding elements in the group. After all groups have completed the operation, 32 elements are obtained. Finally, add these 32 elements together to obtain the final C(i,j). Assume that after matrix A is compressed, it becomes cA with a dimension of m (n / 32); after matrix B_rev is compressed, it becomes cB_rev with a dimension of p (n / 32), then C(i,j) is expressed by the formula:
[0092]
[0093] Among them, cA(i,kb,b) represents the b-th bit of the ikb-th element of matrix cA. The corresponding relationship between cA and A is:
[0094] , where .
[0095] Figure 3 Shows the implementation process schematic diagram of a fast matrix multiplication calculation method according to an embodiment of the present disclosure Figure 3 , Figure 4 Shows the overall implementation scheme schematic diagram of matrix multiplication in the embodiments of the present disclosure, Figure 4 Corresponding to the implementation process in Figure 3 , as Figure 3 Shown, the implementation process of a fast matrix multiplication calculation method according to an embodiment of the present disclosure includes the following steps:
[0096] Step 301, transpose matrix B to obtain B_rev.
[0097] In the embodiments of the present disclosure, before transposing matrix B, it is necessary to first configure the interfaces for data input and output, specifically: Configure the interfaces: Input the dimensions of matrix A: the number of rows row_A of A, and the number of columns col_A of A; Input the dimensions of matrix B: the number of rows row_B of B, and the number of columns col_B of B. According to the operation rules of matrix multiplication, col_A = row_B is required; Data input interface: 32 bits, that is, input one 32-bit data per clock cycle. Input the data of matrix A, with a total of row_A col_A data. Input the data of matrix B, with a total of row_B col_B data. Data output interface: 32-bit, that is, one 32-bit data is output per clock cycle; output the data of matrix C, and a total of row_A data are output col_B data.
[0098] In the embodiments of the present disclosure, the dimensions of matrix A and matrix B are received, and it is determined that if col_A and row_B are not equal, an error is reported. If the amount of data in the input matrix is not equal to the corresponding dimension, an error is reported.
[0099] In the embodiments of the present disclosure, the matrix data can be stored in a random access memory (RAM). Specifically, three RAMs store matrix A, matrix B, and matrix C, namely: RAM_A, RAM_B, and RAM_C, with a width of 32 bits. That is, one 32-bit data can be stored at one address.
[0100] In the embodiments of the present disclosure, since the default storage format is row-major, it is necessary to first transpose matrix B to improve the cache hit rate when performing matrix multiplication on matrix A and matrix B.
[0101] Step 302, compress matrix A and matrix B_rev according to the processor bit width to obtain matrix cA and cB_rev.
[0102] In the embodiments of the present disclosure, first, the processor bit width is determined, denoted as size_bint; the sizeof keyword can be used to measure the actual number of bytes of a variable, and multiplying by 8 gives the actual bit width. Here, it is assumed that size_bint is 32; second, memory space is allocated to construct the compressed matrix. Among them, it is necessary to first read the memory space occupied by matrix A and B, and calculate the memory space required for the compressed matrices cA and cB_rev according to the actual processor bit width and the memory space occupied by matrix A and B. The number of rows of the compressed matrix is the same as that of the original matrix, and the number of columns of the compressed matrix is 1 / 32 of the original matrix. After the matrix is compressed, one number in the matrix represents 32 numbers in the original matrix.
[0103] In the embodiments of the present disclosure, Figure 5 shows a schematic diagram of the specific implementation scheme of matrix compression in the embodiments of the present disclosure. The specific operations include: the corresponding relationship between the matrix elements before and after compression is that the column index of the matrix elements before compression is divided into two parts, where the last 5 bits (log 2 size_bint = log 2 32 = 5) is used as the bit index of the matrix element after compression, representing which bit of the matrix element before compression in the corresponding element of the matrix after compression. The remaining part of the column index of the matrix before compression can be used as the column index of the matrix element after compression, representing which elements of the matrix before compression are included in the matrix element after compression; the corresponding relationship of the matrix elements can be represented by the following C code:
[0104] cop_mat_a[i cop_c + j / size_bint] += Bmat_A[i ca_rb + j] << (j % size_bint)
[0105] Among them, cop_mat_a represents the compressed matrix A, and Bmat_A represents the matrix A before compression. i is the row index, j is the column index, ca_rb represents the number of columns of matrix A and the number of rows of matrix B, and cop_c is equal to ca_rb / size_bint. This line of code means that the corresponding elements of matrix Bmat_A are first left-shifted and then added to cop_mat_a.
[0106] In the embodiments of the present disclosure, the compressed matrix A data is sent to RAM_A, and the compressed / transposed matrix B data is sent to RAM_B.
[0107] In the embodiments of the present disclosure, the valid bits in the input data are intercepted and stored at the corresponding bit positions. The specific method is as follows: The storage space at address 1 in RAM_A is regarded as a 32-bit cache space for storing the compressed data. Only the lowest bit in the input data is valid. The lowest bit is intercepted and placed at the corresponding position in the 32-bit cache space. That is, the valid bit of the first data is stored in the 1st bit of the cache space, the valid bit of the second data is placed in the 2nd bit of the cache space, and so on. After the storage space at address 1 in RAM_A is full, it turns to the next address of the RAM (address 2), regarded as a new 32-bit cache space, to store the next 32 data. Thus, after the input of matrix A is completed, RAM_A will store the compressed data of matrix A. The valid bit of the 32-bit data is intercepted, and this 1-bit data is stored in RAM_B. After the input of matrix B is completed, what is stored in RAM_B is the compressed and transposed data of matrix B.
[0108] Step 303, calculate the multiplication of the compressed matrix.
[0109] In the embodiments of the present disclosure, the data at the corresponding addresses in RAM_A and RAM_B are read out and an AND operation is performed according to the corresponding bit positions.
[0110] In the embodiments of the present disclosure, at this time, the corresponding rows of the cA matrix and the cB_rev matrix should be selected, and the finite field multiplication and addition operations are performed according to matrix multiplication. After the calculation is completed, the matrix element cC(i,j) is obtained. Specifically: take the i-th row of the matrix cA and the j-th row of the matrix cB_rev, and select the corresponding elements in each row to perform a bitwise AND operation (the AND operation is equivalent to the multiplication operation in the gf(2) finite field). Note that here cA and cB_rev are compressed matrices, and each element in these two matrices represents 32 elements of the original matrix. Since the original two matrices follow the same compression scheme, the bitwise AND operation of the compressed matrix elements will not change the corresponding relationship of the original matrix elements.
[0111] Step 304, use finite field addition to add up each bit of the matrix element cC(i,j).
[0112] In the embodiments of the present disclosure, there are two types of exclusive OR operations involved. One is to perform an exclusive OR operation on the corresponding bit positions, and the other is to perform an exclusive OR operation on 32 bits together. Based on the results of the above exclusive OR operations, the new data is added to the old data until the number of accumulation times reaches row_A / 32, and then the accumulated result is subjected to an exclusive OR operation on 32 bits together. The operation result is 1 bit, and the operation result is stored in the corresponding bit position in RAM_C.
[0113] In the embodiments of the present disclosure, the elements after the bitwise AND operation are added up. In the finite field, the addition operation is equivalent to the exclusive OR operation. That is, the data after the AND operation is subjected to a bitwise exclusive OR operation again. In the actual implementation process, the initial value of the exclusive OR operation is assigned 0, and each time a result of the AND operation is obtained, it is added to the result of the previous exclusive OR operation using the exclusive OR operation.
[0114] In the embodiments of the present disclosure, the calculated result has 32 bits, which can be understood as dividing the result of the bitwise AND operation of the original matrix into 32 groups, performing an exclusive OR operation within each group, and finally obtaining 32 numbers. Finally, these 32 numbers need to be subjected to an exclusive OR operation again to finally obtain a 1-bit calculation result. This 1 bit is the actual C(i,j). In the actual implementation process, the initial value of the exclusive OR operation is assigned 0, and each time a result of the AND operation is obtained, it is added to the result of the previous exclusive OR operation using the exclusive OR operation.
[0115] In the embodiments of the present disclosure, after both matrix A and matrix B are input, the matrix calculation is started. The data in RAM_A and RAM_B is read out, and the functions of multiplying corresponding elements and accumulating by rows are executed, and then the operation result is stored in RAM_C.
[0116] Figure 6 The figure shows a schematic diagram of the specific implementation scheme for calculating matrix element grouping in the embodiments of the present disclosure, including the relevant content of the above steps 303 and 304.
[0117] In an embodiment of the present disclosure, preferably, if the matrix dimension does not exceed the RAM storage limit, hardware is used to perform matrix multiplication; if it exceeds, an error is reported, indicating that the hardware cannot perform the operation, and software calculation is used.
[0118] In an embodiment of the present disclosure, a calculation method for matrix multiplication over the finite field gf(2) is proposed, which converts the addition, subtraction, multiplication, and division operations in the rational number field into exclusive OR and operations in the finite field, greatly improving the operation performance; the data in a row of the matrix is divided into groups of 32 each, and each group of data is compressed into one data in the computer. One bit operation of the processor can complete the operation of 32 data in the matrix, thus making full use of the processor resources and greatly improving the operation performance; at the same time, it includes an implementation manner of grouped calculation, and there is no need to restore the compressed matrix.
[0119] Figure 7 Shows a schematic diagram of a fast matrix multiplication calculation device according to an embodiment of the present disclosure Figure 1 , as Figure 7 shown, a fast matrix multiplication calculation device according to an embodiment of the present disclosure includes:
[0120] A data input / output module 701, configured to respectively determine a first matrix and a third matrix based on different storage locations;
[0121] The data input / output module 701 is further configured to determine a first data volume based on the memory sizes occupied by the first matrix and the second matrix; calculate the memory required for the fourth matrix and the fifth matrix based on the actual bit width and the first data volume, determine a second data volume based on the required memory, and allocate memory spaces for the fourth matrix and the fifth matrix based on the second data volume.
[0122] A matrix compression module 702, configured to obtain the actual bit width of the processor, and compress the first matrix and the third matrix based on the actual bit width to obtain a fourth matrix and a fifth matrix;
[0123] The matrix compression module 702 is further configured to divide the column indexes of the first matrix into a first bit index and a first other column index; read the first bit index therein based on the column indexes of the first matrix, and calculate the bit index of the compressed fourth matrix based on the first bit index and the actual bit width; read the first other column index therein based on the column indexes of the first matrix, and calculate the other column index of the compressed fourth matrix based on the first other column index and the actual bit width; construct the fourth matrix based on the bit index and the other column index of the fourth matrix; divide the indexes of the third matrix into a third bit index and a third other index; read the third bit index therein based on the column indexes of the third matrix, and calculate the bit index of the compressed fifth matrix based on the third bit index and the actual bit width; read the third other column index therein based on the column indexes of the third matrix, and calculate the other column index of the compressed fifth matrix based on the third other column index and the actual bit width; construct the fifth matrix based on the bit index and the other column index of the fifth matrix.
[0124] The matrix compression module 702 and the matrix transpose module 705 form a matrix recombination module 709.
[0125] The matrix operation module 703 performs a finite field multiplication and addition operation on the fourth matrix and the fifth matrix using matrix multiplication to obtain a seventh matrix.
[0126] The matrix operation module 703 is further configured to perform a finite field multiplication and addition operation on the fourth matrix and the fifth matrix using matrix multiplication to obtain a sixth matrix;
[0127] The parameter configuration module 704 is configured to respectively receive a first matrix and a second matrix based on different storage locations; respectively read the storage formats of the first matrix and the second matrix; determine the dimension information of the first matrix and the second matrix based on the storage formats, and judge whether the first matrix and the second matrix conform to the matrix multiplication operation rule based on the dimension information; if they do not conform to the matrix multiplication operation rule, generate an error message;
[0128] The matrix transpose module 705 is configured to transpose the storage format of the second matrix based on the matrix multiplication rule to obtain a third matrix;
[0129] The processor bit width calculation module 706 is configured to measure the actual number of bytes of a single data based on a first keyword, and calculate the actual bit width of the processor based on the actual number of bytes of the single data.
[0130] The matrix operation module 703 includes: a matrix accumulation module 707, configured to add each bit of data in the sixth matrix based on finite field addition to obtain a seventh matrix.
[0131] The matrix accumulation module 707 is further configured to perform an exclusive OR operation on the elements after the bitwise AND operation to obtain the sixth matrix.
[0132] The matrix operation module 703 further includes: a matrix multiplication module 708, configured to select corresponding rows from the fourth matrix and the fifth matrix, and perform a bitwise AND operation on the selected corresponding elements in the corresponding rows.
[0133] Figure 8 Fig. shows a schematic diagram of a matrix multiplication fast calculation device according to an embodiment of the present disclosure Figure 2 , such as Figure 8 As shown, a schematic diagram of a matrix fast calculation device according to an embodiment of the present disclosure includes:
[0134] A data input / output module 701, configured to receive matrix data, count the data volume, and transfer the input data to a matrix reorganization module 709. After receiving the operation completion signal sent by the matrix operation module 703, the data in the RAM_C module 712 is sent out.
[0135] A parameter configuration module 704, configured to receive the dimensions of matrix A and matrix B, and report an error if col_A and row_B are not equal. If the data volume of the input matrix is not equal to the corresponding dimension, an error is reported.
[0136] The parameter configuration module 704 can further transfer the matrix dimensions to the matrix reorganization module 709 and the matrix operation module 703.
[0137] A matrix reorganization module 709: Among them, the matrix reorganization module 709 is composed of a matrix compression module 702 and a matrix transpose module 705. If the input matrix is matrix A, the matrix compression function is executed, and the compressed data is sent to the RAM_A module 710. If the input matrix is matrix B, the matrix compression and matrix transpose functions are executed, and the compressed / transposed data is sent to the RAM_B module 711.
[0138] The matrix operation module 703 is configured to start matrix calculation after both matrix A and matrix B are input. The matrix calculation is composed of a matrix multiplication module 708 and a matrix accumulation module 707. The data in the RAM_A module 710 and the RAM_B module 711 is read out, and the functions of multiplying corresponding elements and accumulating by row are executed, and then the operation result is stored in the RAM_C module 712.
[0139] The RAM_A module 710 is configured to store matrix A, with a width of 32 bits. That is, one 32-bit data can be stored at one address.
[0140] The RAM_B module 711 is used to store matrix B, with a width of 32 bits. That is, one 32-bit data can be stored at one address.
[0141] The RAM_C module 712 is used to store matrix C, with a width of 32 bits. That is, one 32-bit data can be stored at one address.
[0142] In the embodiments of the present disclosure, when hardware performs matrix multiplication, parameter configuration needs to be performed first, that is: Configuration interface: Input the dimensions of matrix A: the number of rows row_A of A, and the number of columns col_A of A; input the dimensions of matrix B: the number of rows row_B of B, and the number of columns col_B of B. According to the operation rules of matrix multiplication, col_A = row_B is required; Data input interface: 32 bits, that is, one 32-bit data is input per clock cycle. For the data of input matrix A, a total of row_A col_A data are input. For the data of input matrix B, a total of row_B col_B data are input. Data output interface: 32 bits, that is, one 32-bit data is output per clock cycle; for the data of output matrix C, a total of row_A col_B data are output.
[0143] In the embodiments of the present disclosure, after the hardware completes parameter configuration, relevant data is input, that is, matrix A is compressed and stored, and matrix B is transposed after compression and then stored. Specifically: The matrix recombination module 709 with only compression function: intercepts the valid bits in the input data and stores them in the corresponding bit positions. The specific method is: regard the storage space at address 1 in the RAM_A module 710 as a 32-bit cache space for storing the compressed data. Only the lowest bit in the input data is valid. The lowest bit is intercepted and placed in the corresponding position in the 32-bit cache space. That is, the valid bit of the first data is stored in the 1st bit of the cache space, the valid bit of the second data is placed in the 2nd bit of the cache space, and so on. After the storage space at address 1 in the RAM_A module 710 is full, it turns to the next address of the RAM (address 2), regarded as a new 32-bit cache space, to store the next 32 data. Thus, after matrix A is input, the RAM_A module 710 will store the compressed data of matrix A.
[0144] In the embodiments of the present disclosure, the compression and transposition function of the matrix recombination module 709 is used to process matrix B. Specifically: intercept the valid bits of 32-bit data, and store this 1-bit data in the RAM_B module 711. The position stored in the RAM_B module 711 is calculated by the matrix transposition module 705. After matrix B is input, the data stored in the RAM_B module 711 is the compressed and transposed data of matrix B.
[0145] The matrix multiplication module 708 is also used to read out the data at the corresponding addresses in the RAM_A module 710 and the RAM_B module 711, perform an AND operation on the corresponding bits, and send the obtained result to the matrix accumulation module 707. Specifically: take the i-th row of the matrix cA and the j-th row of the matrix cB_rev, and select the corresponding elements in each row to perform a bitwise AND operation (the AND operation is equivalent to the multiplication operation in the finite field gf(2)). At this time, cA and cB_rev are compressed matrices, and each element in these two matrices represents 32 elements of the original matrix. Since the original two matrices follow the same compression scheme, the bitwise AND operation of the compressed matrix elements will not change the corresponding relationship of the original matrix elements.
[0146] The matrix accumulation module 707 is also used to perform two addition functions, one is to perform an exclusive OR operation on the corresponding bits, and the other is to perform an exclusive OR operation on 32 bits together. Receive the data from the matrix multiplication module 708 and add the new data to the old data until the accumulation count reaches row_A / 32, and then perform an exclusive OR operation on the 32 bits of the accumulated result. The operation result is 1 bit, and the operation result is stored in the corresponding bit of the RAM_C module 712. Specifically: accumulate the elements after the bitwise AND operation. In the finite field, the addition operation is equivalent to the exclusive OR operation. That is, perform a bitwise exclusive OR operation on the data after the AND operation again. Assign the initial value of the exclusive OR operation to 0, and each time a result of the AND operation is obtained, add it to the result of the previous exclusive OR operation using the exclusive OR operation. The accumulated result after the above has 32 bits, which can be understood as dividing the result of the bitwise AND operation of the original matrix into 32 groups, performing an exclusive OR operation within each group, and finally obtaining 32 numbers. Finally, it is also necessary to perform an exclusive OR operation on these 32 numbers again to finally obtain a 1-bit calculation result. This 1 bit is the actual C(i,j). In the actual implementation process, assign the initial value of the exclusive OR operation to 0, and each time a result of the AND operation is obtained, add it to the result of the previous exclusive OR operation using the exclusive OR operation.
[0147] In the embodiment of the present disclosure, after the data input is completed, the internal calculation matrix C is stored and then the data is output.
[0148] In an exemplary embodiment, the data input / output module 701, matrix compression module 702, matrix operation module 703, parameter configuration module 704, matrix transpose module 705, processor bit-width calculation module 706, matrix accumulation module 707, matrix multiplication module 708, matrix recombination module 709, RAM_A module 710, RAM_B module 711, and RAM_C module 712, etc. can be implemented by one or more central processing units (CPUs, Central Processing Unit), graphics processing units (GPUs, Graphics Processing Unit), application specific integrated circuits (ASICs, Application Specific Integrated Circuit), DSPs, programmable logic devices (PLDs, Programmable Logic Device), complex programmable logic devices (CPLDs, Complex Programmable Logic Device), field-programmable gate arrays (FPGAs, Field-Programmable Gate Array), general-purpose processors, controllers, microcontrollers (MCUs, Micro Controller Unit), microprocessors (Microprocessor), or other electronic components.
[0149] Regarding the device in the above embodiment, the specific manner in which each module and unit performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0150] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0151] Figure 9 A schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure is shown. The electronic device 800 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 800 can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] As Figure 9As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the random access memory (RAM) 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the read-only memory (ROM) 802, and the random access memory (RAM) 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0153] Multiple components in the electronic device 800 are connected to the input / output (I / O) interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0154] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units 801 running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as a fast matrix multiplication calculation method. For example, in some embodiments, a fast matrix multiplication calculation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the read-only memory (ROM) 802 and / or the communication unit 809. When the computer program is loaded into the random access memory (RAM) 803 and executed by the computing unit 801, one or more steps of the fast matrix multiplication calculation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute a fast matrix multiplication calculation method in any other appropriate manner (e.g., by means of firmware).
[0155] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0156] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0157] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0159] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0160] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.
[0161] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0162] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of this disclosure, "a plurality" means two or more unless otherwise specifically defined.
[0163] As described above, it is only the specific implementation manner of the present disclosure. However, the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claimed rights.
Claims
1. A fast matrix multiplication calculation method, characterized in that: The method comprises: Based on different storage locations, respectively determine a first matrix and a third matrix; Acquire an actual bit width of a processor, and compress the first matrix and the third matrix based on the actual bit width to obtain a fourth matrix and a fifth matrix; Performing a finite field multiplication and addition operation on the fourth matrix and the fifth matrix by using matrix multiplication to obtain a seventh matrix; The compressing the first matrix and the third matrix based on the actual bit width includes: dividing the column index of the first matrix into a first bit index and a first other column index; Reading a first bit index of the first matrix based on a column index thereof, and calculating a bit index of a compressed fourth matrix based on the first bit index and the actual bit width; Reading a first other column index therein based on the column index of the first matrix, and calculating other column indexes of a compressed fourth matrix based on the first other column index and the actual bit width; Constructing the fourth matrix based on the bit index of the fourth matrix and other column indexes of the fourth matrix; Dividing the index of the third matrix into a third bit index and a third other index; Reading a third bit index of the third matrix based on the column index thereof, and calculating a bit index of a compressed fifth matrix based on the third bit index and the actual bit width; Reading a third other column index therein based on the column index of the third matrix, and calculating other column indexes of the compressed fifth matrix based on the third other column index and the actual bit width; The fifth matrix is constructed based on the bit index of the fifth matrix and other column indexes of the fifth matrix.
2. The method according to claim 1, characterized in that The determining the first matrix and the third matrix respectively based on different storage locations includes: Based on different storage locations, respectively determining the first matrix and the second matrix; Reading storage formats of the first matrix and the second matrix respectively; Determine dimension information of the first matrix and the second matrix based on the storage format, and determine whether the first matrix and the second matrix comply with a matrix multiplication operation rule based on the dimension information; If the matrix multiplication operation rule is not met, an error message is generated; If the matrix multiplication operation rule is met, the storage format of the second matrix is transposed based on the matrix multiplication rule to obtain a third matrix.
3. The method according to claim 1, characterized in that The obtaining the actual bit width of the processor includes: The actual number of bytes of the single data is measured based on the first keyword, and the actual bit width of the processor is calculated based on the actual number of bytes of the single data.
4. The method according to claim 2, characterized in that: The method further comprises: Determine a first data amount based on memory sizes occupied by the first matrix and the second matrix; The memory required for the fourth matrix and the fifth matrix is calculated based on the actual bit width and the first data amount, the second data amount is determined based on the required memory, and memory space is allocated to the fourth matrix and the fifth matrix based on the second data amount.
5. The method according to claim 1, characterized in that The method of performing a finite field multiplication and addition operation on the fourth matrix and the fifth matrix by matrix multiplication to obtain a seventh matrix includes: Performing a finite field multiplication and addition operation on the fourth matrix and the fifth matrix by using matrix multiplication to obtain a sixth matrix; Based on finite field addition, each bit of data in the sixth matrix is added to obtain a seventh matrix.
6. The method according to claim 5, characterized in that The method of performing a finite field multiplication and addition operation on the fourth matrix and the fifth matrix by matrix multiplication to obtain a sixth matrix includes: Selecting corresponding rows in the fourth matrix and the fifth matrix, and performing a bitwise AND operation on the selected corresponding elements in the corresponding rows; An XOR operation is performed on the elements after the bitwise AND operation to obtain the sixth matrix.
7. A fast matrix multiplication calculation device, characterized in that: The device comprises: A data input and output module, used for respectively determining the first matrix and the third matrix based on different storage locations; A matrix compression module, used for obtaining an actual bit width of a processor, and compressing the first matrix and the third matrix based on the actual bit width to obtain a fourth matrix and a fifth matrix; A matrix operation module, used for performing a finite field multiplication and addition operation on the fourth matrix and the fifth matrix by using matrix multiplication to obtain a seventh matrix; The matrix compression module is also used to divide the column index of the first matrix into a first bit index and a first other column index; read the first bit index based on the column index of the first matrix, and calculate the bit index of the compressed fourth matrix based on the first bit index and the actual bit width; read the first other column index based on the column index of the first matrix, and calculate the other column index of the compressed fourth matrix based on the first other column index and the actual bit width; construct the fourth matrix based on the bit index of the fourth matrix and the other column index of the fourth matrix; divide the index of the third matrix into a third bit index and a third other index; read the third bit index based on the column index of the third matrix, and calculate the bit index of the compressed fifth matrix based on the third bit index and the actual bit width; read the third other column index based on the column index of the third matrix, and calculate the other column index of the compressed fifth matrix based on the third other column index and the actual bit width; construct the fifth matrix based on the bit index of the fifth matrix and the other column index of the fifth matrix.
8. The device according to claim 7, characterized in that The device also includes: a parameter configuration module, configured to respectively determine the first matrix and the second matrix based on different storage locations; respectively read the storage formats of the first matrix and the second matrix; determine the dimension information of the first matrix and the second matrix based on the storage formats, and determine whether the first matrix and the second matrix comply with the matrix multiplication operation rules based on the dimension information; and generate an error message if they do not comply with the matrix multiplication operation rules; A matrix transposition module, used for transposing the storage format of the second matrix based on a matrix multiplication rule to obtain a third matrix; The data input and output module is also used to determine a first data amount based on the memory size occupied by the first matrix and the second matrix; calculate the memory required for the fourth matrix and the fifth matrix based on the actual bit width and the first data amount, determine the second data amount based on the required memory, and allocate memory space for the fourth matrix and the fifth matrix based on the second data amount.
9. The device according to claim 7, characterized in that The matrix compression module also includes: The processor bit width calculation module is used to measure the actual number of bytes of a single data based on the first keyword, and calculate the actual bit width of the processor based on the actual number of bytes of the single data.
10. The device according to claim 7, characterized in that The matrix operation module is further used to perform a finite field multiplication and addition operation on the fourth matrix and the fifth matrix using matrix multiplication to obtain a sixth matrix; The matrix operation module includes: a matrix accumulation module, which is used to add each bit of data in the sixth matrix based on finite field addition to obtain a seventh matrix.
11. The device according to claim 10, characterized in that The matrix operation module also includes: a matrix multiplication module, which is used to select corresponding rows in the fourth matrix and the fifth matrix, and perform a bitwise AND operation on the selected corresponding elements in the corresponding rows; The matrix accumulation module is further used to perform an XOR operation on the elements after the bitwise AND operation to obtain the sixth matrix.
12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Matrix multiplier
CN111859273A
Graphic processor, matrix multiplication task processing method and device and storage medium
CN115880132A
Memory allocation method, accelerator and storage medium
CN117349190A