High-performance method and device for left-right multiplication of three-valued matrix and arbitrary matrix
Through the hardware architecture of chunked parallel iteration and the design of preprocessing addition tree modules, the problem of waste and low efficiency of computing resources for the left and right multiplication of three-value matrix and any matrix is solved, and an efficient computing process is realized.
Patent Information
- Application Number
- CN202510435543.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
When performing left-right multiplication of three-value matrix and any matrix, the prior art requires a large demand for computing resources and is inefficient, and fails to fully utilize the special properties of the three-value matrix, resulting in waste of hardware resources and bottlenecks in computing speed.
Using a block, parallel and iterative hardware architecture, the calculation is performed through the block matrix multiplication processing unit and the vector multiplication processing unit, and the element multiplication and accumulation are used to multiply and accumulate elements to avoid the use of traditional multipliers.
It significantly reduces the demand for hardware resources, improves computing efficiency and speed, simplifies the hardware structure, and shortens the delay of critical paths.
Smart Images

Figure CN120277312A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computers and integrated circuits, and particularly to a high-performance method and device for left and right multiplication of a ternary matrix and an arbitrary matrix. Background Art
[0002] Matrix multiplication is a basic operation in various technical fields and plays a crucial role in fields such as artificial intelligence and lattice cryptography. In artificial intelligence, matrix multiplication supports the inference and training of neural networks and is a core computing operation; while in lattice cryptography algorithms, matrix multiplication operations play a key role in the encryption and decryption processes. In practical applications, a specific form of matrix multiplication is often encountered: the multiplication operation between an arbitrary matrix and a ternary matrix. Especially for the Scloud+ algorithm in lattice cryptography, the left and right multiplication operations between an arbitrary matrix and a ternary matrix have become a performance bottleneck.
[0003] Generally speaking, when performing left and right multiplication operations of a ternary matrix and an arbitrary matrix, when the matrix dimension is high and the element values are large, the demand for computing resources increases significantly; at the same time, the inconsistency of data access patterns also increases the complexity of data transmission, further restricting the improvement of algorithm efficiency. Currently, common matrix multiplication schemes mostly perform block processing on matrices and then use fully connected matrix multipliers to complete the computing tasks. However, this method fails to fully utilize the special properties of the ternary matrix, resulting in waste of a large amount of multiplier resources. Although this scheme can complete the necessary calculations, there are still obvious bottlenecks in terms of area utilization and computing speed, and it is difficult to achieve efficient performance improvement. Summary of the Invention
[0004] The present application provides a high-performance method and device for left and right multiplication of a ternary matrix and an arbitrary matrix, which is particularly suitable for the case of a relatively large matrix scale, to solve the problems of large computational amount and low efficiency when calculating left and right multiplication of a large-scale arbitrary matrix and a ternary matrix in the prior art.
[0005] In a first aspect, the present application provides a high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix, including:
[0006] A block matrix multiplication processing unit configured to perform multiplication processing on an arbitrary square matrix and a ternary square matrix; the arbitrary square matrix is a block of an arbitrary matrix, and the ternary square matrix is a block of a ternary matrix;
[0007] The block matrix multiplication processing unit includes a plurality of vector multiplication processing units; the vector multiplication processing unit is configured to perform multiplication processing on the row elements of the arbitrary square matrix and the column elements of the ternary square matrix to generate vector products;
[0008] The block matrix multiplication processing unit is further configured to generate a product matrix of the arbitrary square matrix and the ternary square matrix according to a plurality of vector products obtained by the plurality of vector multiplication processing units;
[0009] Obtain a product result matrix of the arbitrary matrix and the ternary matrix according to a plurality of product matrices of a plurality of arbitrary square matrices and a plurality of ternary square matrices;
[0010] A memory, configured to store the product matrix and the product result matrix.
[0011] In some embodiments, the vector multiplication processing unit includes a preprocessing module and an adder tree module;
[0012] The vector multiplication processing unit performs multiplication processing on the row elements of the arbitrary square matrix and the column elements of the ternary square matrix, and is specifically configured to:
[0013] Calculate the element product of the row element and the column element through the preprocessing module;
[0014] Calculate the product sum of the element products through the adder tree module, and mark the product sum as a vector product.
[0015] In some embodiments, the block matrix multiplication processing unit is further configured to:
[0016] Calculate the number of vector products according to the preset dimension;
[0017] Divide the arbitrary square matrix into multiple rows of elements, and divide the ternary square matrix into multiple columns of elements;
[0018] Input the multiple rows of elements and the multiple columns of elements into a plurality of vector multiplication processing units corresponding to the number of vector products; wherein, one row of elements of the arbitrary square matrix and one column of elements of the ternary square matrix are input into each vector multiplication processing unit.
[0019] In some embodiments, the preprocessing module performs calculation of the element product of the row element and the column element, and is specifically configured to:
[0020] Number the row elements according to the arrangement order of the row elements, and number the column elements according to the arrangement order of the column elements;
[0021] Perform product calculation on the row element and the column element with the same number to obtain the element product.
[0022] In some embodiments, the preprocessing module performs product calculation on the row element and the column element with the same number, and is specifically configured to:
[0023] Obtain the target value corresponding to the column element; wherein, the column element of the three-value square matrix is a two-bit binary value; when the column element is 00, the target value corresponding to the column element is 0; when the column element is 01, the target value corresponding to the column element is 1; when the column element is 11, the target value corresponding to the column element is -1;
[0024] Calculate the product of the row element and the target value, and mark the product as the element product of the row element and the column element.
[0025] In some embodiments, the preprocessing module includes a preset number of bit calculation units; the preset number is the number of bits of the row element;
[0026] The bit calculation unit includes an AND gate, a NOT gate, and a multiplexer; the output end of the AND gate is connected to the input end of the multiplexer, and the output end of the NOT gate is connected to the input end of the multiplexer;
[0027] The preprocessing unit performs a product calculation on the row element and the column element with the same number, and is specifically configured to:
[0028] Input one-bit values of the row element into the AND gate and the NOT gate respectively, and input the lowest-bit value of the column element into the AND gate; the column element is a two-bit binary value;
[0029] Input the first processing result of the NOT gate and the second processing result of the AND gate into the multiplexer;
[0030] Detect the highest-bit value of the column element;
[0031] If the highest-bit value is 1, output the first processing result through the multiplexer and mark it as the bit product result corresponding to the one-bit value; if the highest-bit value is 0, output the second processing result through the multiplexer and mark it as the bit product result corresponding to the one-bit value;
[0032] Combine the bit product results corresponding to the values of all bits of the row element to obtain the element product of the row element and the column element.
[0033] In some embodiments, the adder tree module calculates the product sum of the element products, and is specifically configured to:
[0034] Multiple element products of multiple row elements and multiple column elements are input into the adder tree module;
[0035] The addition tree module compresses the multiple element products and the product sum of the previous clock cycle to generate the accumulated sum output of the current clock cycle; the product sum of the previous clock cycle is the vector product obtained when the previous arbitrary square matrix is multiplied by the previous ternary square matrix.
[0036] Mark the accumulated sum output of the current clock cycle as the product sum of the element products.
[0037] In some embodiments, the addition tree module includes a Wallace tree and a carry - look - ahead adder; the Wallace tree includes a plurality of carry - save adders, and the plurality of carry - save adders are used to compress multiple values into two values.
[0038] The addition tree module compresses the multiple element products and the product sum of the previous clock cycle, and is further configured to:
[0039] Obtain the product sum of the previous clock cycle;
[0040] Process the multiple element products and the product sum of the previous clock cycle through the Wallace tree, and fill the least significant bit of the carry output value of the carry - save adder with the most significant bit value of the column element to obtain two outputs of the Wallace tree.
[0041] In some embodiments, after the addition tree module obtains two outputs of the Wallace tree, it is further configured to:
[0042] Use the most significant bit value of the column element as the carry input of the carry - look - ahead adder, and sum the two outputs of the Wallace tree through the carry - look - ahead adder to generate the accumulated sum output of the current clock cycle.
[0043] In a second aspect, the present application further provides a high - performance method for left - right multiplication of a ternary matrix and an arbitrary matrix, and the multiplication processing method is applied to the above - mentioned device; the method includes:
[0044] Perform multiplication and accumulation processing on the row elements of an arbitrary square matrix and the column elements of a ternary square matrix to generate a vector product; the arbitrary square matrix is a block of an arbitrary matrix, and the ternary square matrix is a block of a ternary matrix.
[0045] Generate the product matrix of the arbitrary square matrix and the ternary square matrix according to the obtained multiple vector products.
[0046] Obtain the product result matrix of the arbitrary matrix and the ternary matrix according to the multiple product matrices of multiple arbitrary square matrices and multiple ternary square matrices.
[0047] As can be seen from the above, the present application provides a high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix. The device includes a block matrix multiplication processing unit and a memory. The block matrix multiplication processing unit includes a plurality of parallel vector multiplication processing units, which can calculate the multiplication of a ternary matrix block and an arbitrary matrix block once in one clock cycle. The vector multiplication processing unit can perform multiplication processing and accumulation on a row of elements of an arbitrary square matrix and a column of elements of a ternary square matrix to generate a vector product. The output of the block matrix multiplication processing unit is combined and accumulated to generate the final product result matrix of the arbitrary matrix and the ternary matrix. The memory can store the results of the block matrix multiplication processing unit. By means of a block, iterative, and parallel computing architecture, the hardware structure is simplified, the requirements for hardware resources are significantly reduced, and the computing efficiency is improved, thus making the entire computing process more efficient and economical. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0049] Figure 1 FIG. shows the overall structure diagram of the high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix in an embodiment of the present application;
[0050] Figure 2 FIG. shows the input / output schematic diagram of the block matrix multiplication processing unit in an embodiment of the present application;
[0051] Figure 3 FIG. shows the internal structure schematic diagram of the vector multiplication processing unit in an embodiment of the present application;
[0052] Figure 4 FIG. shows the internal structure schematic diagram of the preprocessing module in an embodiment of the present application;
[0053] Figure 5 FIG. shows the internal structure schematic diagram of the adder tree module in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The embodiments will be described in detail below, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0055] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the following described embodiments, rather than intending to limit the embodiments of this application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0056] In this application, the terms "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0057] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0058] In this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0059] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to this element.
[0060] Matrix multiplication is a fundamental operation widely used in various technical fields, involving multiple fields such as artificial intelligence and lattice cryptography, and is particularly crucial in the processes of neural network inference and encryption and decryption of lattice cryptography algorithms. In the above applications, a common form of matrix multiplication is the multiplication operation of an arbitrary matrix and a ternary matrix. Especially in the lattice cryptography Scloud+ algorithm, the left and right multiplications of an arbitrarily large matrix and a ternary large matrix have become a performance bottleneck. This process has characteristics such as high dimension, large integer type, and inconsistent memory access patterns for left and right multiplications, which poses challenges for software and hardware implementation. Existing hardware implementations use fully connected matrix multipliers to calculate matrix multiplication, without making full use of the unique properties of ternary matrices, consuming a large amount of multiplier resources and resulting in low area utilization efficiency and computing efficiency.
[0061] To solve the above problems, the embodiment of the present application provides a high-performance device for left-right multiplication of a ternary matrix and an arbitrary matrix, which can calculate and process the multiplication of an arbitrary matrix and a ternary matrix. The device significantly reduces the demand for hardware resources during large-scale matrix multiplication operations through a block-based, parallel and iterative hardware architecture. At the same time, the multiplication operation module between vectors is designed to require only a simple preprocessing unit and a tree-type adder to complete, without the participation of a traditional multiplier. This design not only simplifies the hardware structure, but also shortens the critical path delay, significantly improving the calculation speed.
[0062] In the embodiment of the present application, any matrix can be expressed as Each element in any matrix can be a number less than q=2 l Any non-negative integer, the dimension of any matrix is represented as m×n.
[0063] A ternary matrix refers to a matrix in which the elements have only three numerical conditions, that is, the value range of each element of the ternary matrix has only three values {-1, 0, 1}. In the embodiment of the present application, the ternary matrix is also called a secret matrix.
[0064] The multiplication of a ternary matrix S and an arbitrary matrix A can be divided into two categories, including left multiplication and right multiplication. The left multiplication can be expressed as A·S, and the dimension of S is The three-value matrix is expressed as When multiplied on the right, it can be expressed as S'·A, and the dimension of S' is When multiplied on the right, the three-value matrix is expressed as
[0065] It should be noted that, since S'·A=(A T ·S′ T ) T , the case of right-multiplying matrix A can be converted to left-multiplying A T That is, the processing method of right multiplication can be converted into the transposition of left multiplication. In the embodiment of the present application, the method of left multiplication of a ternary matrix and an arbitrary matrix is used as an example for explanation.
[0066] Figure 1 The overall structure diagram of the high-performance device for left-right multiplication of a ternary matrix and an arbitrary matrix in the embodiment of the present application is shown. Figure 1 As shown, a high-performance device for left-right multiplication of a ternary matrix with an arbitrary matrix includes a memory and a block matrix multiplication processing unit.
[0067] The input of the block matrix multiplication processing unit (BMMPE) is a block of an arbitrary matrix A and a ternary matrix S or S', and the block size is a preset dimension b × b. In the embodiment of the present application, the block of an arbitrary matrix is called an arbitrary block matrix, and the block of a ternary matrix is called a ternary block matrix.
[0068] The memory is configured to store the results of the operations of the block matrix multiplication processing unit, including the calculated matrix.
[0069] A block matrix multiplication processing unit (BMMPE) is configured to perform multiplication processing on an arbitrary matrix and a ternary matrix.
[0070] It should be noted that during the process of calculating the multiplication of an arbitrary matrix and a ternary matrix, since the scales of the two matrices may be large, directly multiplying the two matrices will result in a huge amount of calculation and low calculation efficiency. Therefore, in the embodiments of the present application, the arbitrary matrix and the ternary matrix are processed after being partitioned, so as to convert the matrix multiplication with a large scale into the matrix multiplication with a smaller scale after partitioning.
[0071] Among them, b can be 8, that is, the partitioned matrix is a small-scale matrix of 8×8. Any square matrix can be represented as A b ×b , and the ternary square matrix can be represented as S b×b .
[0072] The block matrix multiplication processing unit can process a group of partitioned matrices each time, that is, multiply an arbitrary square matrix and a ternary square matrix. In the embodiments of the present application, the process of multiplying the partitioned matrices by the block matrix multiplication processing unit once is also called a clock cycle.
[0073] If the size of the partitioned square matrix is b×b, then the arbitrary matrix A is divided into blocks, and the ternary matrix S is divided into blocks, then the original large-scale matrix multiplication can be converted into iterations of the multiplication of square matrices.
[0074] In the embodiments of the present application, the matrix inputs of left multiplication and right multiplication can be controlled by a multiplexer. Figure 2 The input / output schematic diagram of the block matrix multiplication processing unit is shown. As Figure 2 shown, when the matrix is subjected to left multiplication processing, the selection signal of the multiplexer is 0, the left matrix input to the BMMPE is the matrix after partitioning the arbitrary matrix A, and the right matrix is the matrix after partitioning the ternary matrix S.
[0075] When the matrix is subjected to right multiplication processing, the input signal of the multiplexer is 1, the left matrix input to the BMMPE is the matrix after partitioning A T after partitioning, and the right input is S′ TThe matrix after partitioning. Therefore, the left input matrix of BMMPE must be a block matrix from an arbitrary matrix, and the right matrix must be a block matrix from a ternary matrix. The right multiplication case can be transformed into the transpose of the left multiplication, thus ensuring the consistency of the memory access patterns for both left and right multiplications.
[0076] In the embodiments of the present application, the block matrix multiplication processing unit can calculate the multiplication process of an arbitrary block matrix and a ternary block matrix. Both the arbitrary block matrix and the ternary block matrix are matrices of size b×b.
[0077] The arbitrary block matrix can be represented as:
[0078]
[0079] where, a i,j represents the element in the i-th row and j-th column of the arbitrary block matrix.
[0080] The ternary block matrix can be represented as:
[0081]
[0082] where, s j,k represents the element in the j-th row and k-th column of the ternary block matrix.
[0083] The arbitrary block matrix and the ternary block matrix are input into the block matrix multiplication processing unit, and the block matrix multiplication processing unit multiplies the arbitrary block matrix and the ternary block matrix.
[0084] When performing the multiplication of two matrices, the vector product of the row vector formed by each row element in the left matrix and the column vector formed by each column element in the right matrix can be calculated respectively. The vector product is the element in the new matrix after the matrix multiplication.
[0085] Therefore, the multiplication of an arbitrary block matrix and a ternary block matrix of size b×b can be split into b 2 vector multiplications, that is, the b row vectors of the arbitrary block matrix are multiplied by the b column vectors of the ternary block matrix respectively, obtaining b 2 vector products, thus forming a new matrix of size b×b after multiplication.
[0086] In the embodiments of the present application, the block matrix multiplication processing unit may include multiple vector multiplication processing units (PE). Each vector multiplication processing unit can process a vector multiplication process. Therefore, the block matrix multiplication processing unit may include b 2 vector multiplication processing units.
[0087] The vector multiplication processing unit performs multiplication and accumulation processing on the row elements of the arbitrary block matrix and the column elements of the ternary block matrix to generate vector products.
[0088] Considering the particularity of the element values of the ternary matrix, in the embodiments of the present application, a multiplier is not used to calculate the multiplication of matrix elements. Instead, a preprocessing unit is used to calculate the multiplication between vector elements of the block matrix respectively.
[0089] A vector formed by the elements of a row of any block matrix includes b elements, and a vector formed by the elements of a column of a ternary block matrix also includes b elements. Therefore, the multiplication process of a row of elements of any block matrix and a column of elements of a ternary block matrix can be split into multiplying an element of any block matrix by an element of a ternary block matrix, and then summing the results of multiplying multiple groups of elements.
[0090] Taking b = 8 as an example, a row of elements of any block matrix can be expressed as (a1, a2, a3, a4, a5, a6, a7, a8), and a column of elements of a ternary block matrix can be expressed as (s1, s2, s3, s4, s5, s6, s7, s8). Then the multiplication of the row elements and the column elements can be expressed as: a1×s1 + a2×s2 + a3×s3 + a4×s4 + a5×s5 + a6×s6 + a7×s7 + a8×s8.
[0091] Each vector multiplication processing unit can execute the above process of multiplying a row of elements by a column of elements.
[0092] After an arbitrary block matrix and a ternary block matrix are input into the block matrix multiplication processing unit, the block matrix multiplication processing unit can calculate the number of vector products according to a preset dimension. The dimensions of both the arbitrary block matrix and the ternary block matrix are b×b. Therefore, b 2 vector multiplications need to be executed, that is, the number of vector products is b 2 .
[0093] The block matrix multiplication processing unit can divide an arbitrary block matrix into multiple rows of elements, and, divide the ternary block matrix into multiple columns of elements.
[0094] The block matrix multiplication processing unit can input the multiple rows of elements and multiple columns of elements into multiple vector multiplication processing units corresponding to the number of vector products. Among them, a row of elements of an arbitrary block matrix and a column of elements of a ternary block matrix are input into each vector multiplication processing unit.
[0095] For example, (1, 1) can represent the first row of elements of an arbitrary block matrix and the first column of elements of a ternary block matrix. Then the multiple rows of elements and multiple columns of elements can have b 2 element combinations: (1, 1), (1, 2)...(1, b)...(b, b). The b 2 element combinations are respectively input into b 2 vector multiplication processing units.
[0096] In the embodiments of the present application, the vector multiplication processing unit includes a preprocessing module and an adder tree module. Among them, the preprocessing module can perform the process of element multiplication, and the adder tree module, also known as the adder tree, can perform the process of summing after element multiplication.
[0097] Figure 3 A schematic diagram of the vector multiplication processing unit in the embodiments of the present application is shown. As Figure 3 shown, after any matrix and the ternary matrix are partitioned, a row of elements of an arbitrary square matrix and a column of elements of the ternary square matrix are input into the preprocessing module. The preprocessing module calculates the element products and inputs them into the adder tree, and the adder tree can calculate the vector multiplication result value, that is, the vector product.
[0098] The vector multiplication processing unit can number the row elements according to the arrangement order of the row elements, and, number the column elements according to the arrangement order of the column elements.
[0099] The preprocessing module can perform product calculation on the row elements and column elements with the same number to obtain the element products.
[0100] Multiple preprocessing modules can be provided in the vector multiplication processing unit, and each preprocessing module can perform product calculation on a group of row elements and column elements. Among them, the number of a row of elements and a column of elements is both b, that is, the vector multiplication processing unit needs to process the multiplication of b groups of elements. Therefore, b preprocessing modules can be provided in the vector multiplication processing unit. The b preprocessing modules can perform the process of element multiplication in parallel to improve the calculation efficiency of matrix multiplication.
[0101] It should be noted that the row elements are the elements of an arbitrary square matrix, which can be elements with l bits, and can be expressed as:
[0102] a i,j =(a [l-1] i,j ,…,a [0] i,j )
[0103] Among them, [l - 1] represents the highest bit of the row element, and [0] represents the lowest bit of the row element.
[0104] The column elements are the elements of the ternary square matrix, which are binary numerical elements with 2 bits, and can be expressed as:
[0105] s j,k =(s [1] j,k ,s [0] j,k )
[0106] Among them, [1] represents the highest bit of the column element, and [0] represents the lowest bit of the column element.
[0107] The 2-bit column element can be converted into three numerical values. Among them, when the column element is 00, the target numerical value corresponding to the column element is 0; when the column element is 01, the target numerical value corresponding to the column element is 1; when the column element is 11, the target numerical value corresponding to the column element is -1.
[0108] The preprocessing module can obtain the target numerical value corresponding to the column element, calculate the product of the row element and the target numerical value, and mark the product as the element product of the row element and the column element.
[0109] In the embodiment of the present application, the product calculation can be performed on the row element and the column element with the same number by the preprocessing module. The preprocessing module processes each bit of the row element by using the value of the column element, including bitwise inversion, bitwise clearing to 0, or maintaining the original value, so as to output the final product result.
[0110] The preprocessing module includes a preset number of bit calculation units, where the preset number is the number of bits l of the row element. Each bit calculation unit can process one bit value of the row element by using the column element.
[0111] The bit calculation unit includes an AND gate, a NOT gate, and a multiplexer. Among them, the output end of the AND gate is connected to the input end of the multiplexer, and the output end of the NOT gate is connected to the input end of the multiplexer.
[0112] Figure 4 Shows a schematic diagram of the preprocessing module. As Figure 4 shown, the preprocessing module includes a plurality of bit calculation units, and each bit calculation unit processes one bit value of the row element.
[0113] When performing product calculation on the row element and the column element with the same number, the control unit can respectively input one bit value of the row element to the AND gate and the NOT gate, and input the lowest bit value of the column element to the AND gate.
[0114] Taking Figure 4 the lowest bit value a i,j of the row element a [0] i,j in [0] i,j as an example, a j,k can be respectively input to the AND gate and the NOT gate. At the same time, the highest bit value s [1] j,k of the column element s [0] j,k is input to the multiplexer for judgment, and the lowest bit value s
[0115] The NOT gate can invert a single-bit value of a row element. In the embodiments of this application, the result output by the NOT gate is referred to as the first processing result. The AND gate can multiply a single-bit value of a row element and the least significant bit value of a column element. In the embodiments of this application, the result output by the AND gate is referred to as the second processing result.
[0116] The first processing result of the NOT gate and the second processing result of the AND gate are input to the multiplexer. At the same time, the most significant bit value of the column element is also input to the multiplexer, and the multiplexer can select the output result according to the most significant bit value of the column element.
[0117] The multiplexer can detect the most significant bit value of the column element.
[0118] If the most significant bit value is 1, the multiplexer outputs the first processing result, which is marked as the product result corresponding to this bit value of the row element, and is referred to as the bit product result in the embodiments of this application.
[0119] If the most significant bit value is 0, the multiplexer outputs the second processing result, which is marked as the bit product result corresponding to this bit value of the row element;
[0120] Multiple bit calculation units respectively process a single-bit value of the row element, so as to obtain the bit product results of all bit values of the row element. The preprocessing module can combine the bit product results corresponding to the values of all bits of the row element to obtain the element product of the row element and the column element.
[0121] Table 1 shows the schematic values of the element products of the row element and the column element.
[0122] Table 1
[0123]
[0124] As shown in Table 1, when the column elements are 00, 01, and 11, the corresponding target values are 0, 1, and -1 respectively. Therefore, the element products of the row element and the column element are 0, a i,j and a i,j bitwise inversion.
[0125] Combining the bit product results corresponding to the values of all bits of the row element to obtain the element product can be expressed as:
[0126]
[0127] After multiplying b row elements and b column elements, b element products can be obtained, and the element product is a value with l bits. These b elements are input to the subsequent adder tree module to calculate the product sum of the b element products.
[0128] The addition tree module can accumulate the product of b elements and the product sum of the previous clock cycle to generate the accumulated sum output of the current clock cycle, and can mark the accumulated sum output of the current clock cycle as the product sum of the element products. Among them, the product sum of the previous clock cycle is the vector product obtained when the previous arbitrary square matrix is multiplied by the previous ternary square matrix, that is, the vector product calculated in the same PE unit in the previous clock cycle. The reason for adding the product sum of the previous cycle is that in matrix block multiplication, the large matrix needs to be divided into blocks and multiplied separately, and each block in the final result matrix is the accumulation of the product matrices of multiple square matrices, which is an iterative process.
[0129] In the embodiment of the present application, the addition tree module includes a Wallace tree and a ripple-carry adder (RCA). Among them, the Wallace tree includes several carry-save adders (CSA), and several CSAs are used to compress b + 1 input values into 2 values, and the two compressed values are then input into the RCA for bit-by-bit summation.
[0130] This is because if the RCA is directly used to implement the accumulation operation of multiple l-bit elements, it will lead to an overly long critical path, resulting in significant latency. To avoid this situation, in the embodiment of the present application, a Wallace tree composed of multiple CSAs is used to compress b + 1 inputs into 2 inputs, and finally the RCA is used to calculate the final result of the two l-bit inputs, thus greatly reducing the critical path.
[0131] In addition, it should be noted that when the column element is 11, the corresponding target value is -1, and the product result of the row element and the column element should be -a i,j . In a computer, negative numbers are usually represented in two's complement. The two's complement of a number is obtained by taking the bitwise complement (i.e., flipping each bit, changing 0 to 1 and 1 to 0) of the binary representation of its absolute value and then adding 1. That is, -a i,j is equivalent to taking the bitwise complement of a i,j and then adding 1, which can be expressed as:
[0132]
[0133] However, as mentioned above, in the preprocessing module, only the bitwise complement operation is performed on a i,j using a NOT gate, and the addition of 1 is missing. Therefore, it is necessary to perform the addition of 1 operation on the bitwise complemented row element after preprocessing in the addition tree module, which can be achieved by adding the highest bit s[1] of b s j,k in the addition tree module 0 , s[1] 1 … s[1] b-1It is achieved because only the most significant bit of the column element corresponding to the row element that is bitwise inverted after preprocessing is 1, and it is 0 in other cases. The two outputs of the CSA in the adder tree are carry and sum. For the carry output, the least significant bit (LSB) is zero. Therefore, the increment operation can be transformed into filling s in the LSB of the carry output of the CSA. j,k The 1-bit most significant bit value of is completed. The purpose of doing this is to avoid introducing an adder in the preprocessing module, causing additional overhead.
[0134] That is to say, in addition to compressing the product sum of multiple elements and the product sum of the previous clock cycle, the Wallace tree in the adder tree also fills the most significant bit value of the column element in the least significant bit of the carry output value of the CSA to obtain the final two outputs.
[0135] Taking b = 8 as an example, the adder tree module includes 7 CSAs and 1 RCA. Figure 5 The schematic diagram of the adder tree module in some embodiments is shown. As Figure 5 shown, 7 CSAs form a Wallace tree, which can compress the product of b elements and the accumulated sum of the previous cycle to obtain two outputs. At the same time, the most significant bit value of the column element is filled in the carry output of each CSA, that is, the value of the least significant bit of the carry output value is changed to the most significant bit value of the column element. The two outputs of the Wallace tree can be input into the RCA, and the RCA performs an accumulation calculation. Here, note that since we have the most significant bit values of 8 column elements, but there are only 7 carry outputs of the CSA, the most significant bit value of the remaining 1 column element can be used as the carry input of the ripple-carry adder and input into the RCA for summation. The adder tree module also includes a register D, which can store the accumulated sum of the previous cycle.
[0136] The adder tree module can sum the product of b elements and the product sum of the previous cycle to generate the accumulated sum output of the current clock cycle. This accumulated sum output is the vector product of the row elements of any square matrix and the column elements of the ternary square matrix.
[0137] It should also be noted that in order to ensure the consistency of the number of bits during the accumulation process, in the application of the cryptographic algorithm, the output of each sum of the CSA and RCA needs to be modulo q = 2 l operation, which can be completed by directly omitting the most significant bit (MSB) of each adder (including CSA and RCA). For applications in neural networks, the quantization operation can be completed by the least significant bit (LSB) of each adder (including CSA and RCA). These operations do not require additional calculation after the adder tree to reduce the consumption of critical paths and hardware resources.
[0138] The present application also provides a high-performance method for left and right multiplication of a ternary matrix and an arbitrary matrix, which is applied to the above-mentioned device; the method is as follows: divide the ternary matrix and the arbitrary matrix into b×b squares, and then iteratively perform square matrix multiplication calculations. According to the accumulation and combination of multiple product matrices of multiple arbitrary square matrices and multiple ternary square matrices, the product result matrix of the arbitrary matrix and the ternary matrix is obtained. By selecting an appropriate parameter b for block division, the optimal number of iterations can be achieved.
[0139] As can be seen from the foregoing, the arbitrary matrix A includes blocks of arbitrary square matrices, and the ternary matrix S includes blocks of ternary square matrices. Therefore, product matrices can be obtained. These product matrices can be accumulated to obtain accumulation matrices.
[0140] Since the block matrix multiplication processing unit can only calculate one product matrix per cycle, therefore the accumulation of product matrices in
[0141] cycles can obtain the accumulation matrix. After obtaining an accumulation matrix, the data in the register can be cleared, and the next accumulation matrix can be obtained. Multiple accumulation matrices can form the final result matrix, that is, the multiplication processing result of the arbitrary matrix and the ternary matrix.
[0142] Table 2 shows the content of the block matrix multiplication algorithm.
[0142] Table 2
[0143]
[0144] In this specification, the same or similar parts among the various embodiments can be referred to each other, and will not be elaborated here.
[0145] This embodiment has the following advantages:
[0146] 1. A general block matrix multiplication method based on a processing unit: A general block matrix multiplication method based on a processing unit is proposed. The method divides the matrix into b×b squares, and then iteratively performs square matrix multiplication calculations. Among them, the multiplication of b×b squares is calculated in parallel by multiple processing units PE. One PE calculates the multiplication of a b-value vector. Then, a total of b 2 PEs are required for each square matrix multiplication calculation. By selecting an appropriate parameter b for block division, the optimal number of iterations can be achieved.
[0147] 2. High-performance hardware architecture design: A high-performance hardware architecture for the left and right multiplication of ternary matrices and arbitrary matrices is designed according to this method. This architecture is based on the above-mentioned block and iteration methods, combined with a preprocessing unit and compact negative number processing, saving circuit area and simplifying the hardware structure. At the same time, through the parallel design of the underlying processing unit and the use of an adder tree to shorten the critical path, the computing efficiency is improved as much as possible, making the entire computing process more efficient and economical. Specifically:
[0148] (1) During the multiply-accumulate calculation process of the PE, the multiplication operation can be completed by a preprocessing module composed of AND gates, inverters, and multiplexers, thus reducing the demand for hardware resources.
[0149] (2) During the accumulation process, this design uses an adder tree structure composed of a Wallace tree and an RCA, so that the critical path only has full adders and 1 carry-ripple adder, reducing the delay of the critical path.
[0150] (3) The negative number operation is equivalent to the complement operation. The complement operation can be split into two steps: inversion and addition by one. These two steps are performed separately in the preprocessing module and the adder tree module. In the adder tree, the addition by one operation can be performed by directly filling the highest bit of the column element, avoiding the introduction of an adder in the preprocessing module, thereby reducing the hardware overhead in negative number processing.
[0151] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments of the present invention.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0153] For the sake of convenience in explanation, the above description has been made in connection with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix, characterized in that Including: A block matrix multiplication processing unit configured to perform multiplication processing on any square matrix and a ternary square matrix; The any square matrix is a block of any matrix, and the ternary square matrix is a block of a ternary matrix; The block matrix multiplication processing unit includes a plurality of vector multiplication processing units; the vector multiplication processing unit is configured to perform multiplication processing on the row elements of the any square matrix and the column elements of the ternary square matrix to generate a vector product; The block matrix multiplication processing unit is further configured to generate a product matrix of the any square matrix and the ternary square matrix according to a plurality of vector products obtained by the plurality of vector multiplication processing units; Obtain a product result matrix of the any matrix and the ternary matrix according to a plurality of product matrices of a plurality of any square matrices and a plurality of ternary square matrices; A memory configured to store the product matrix and the product result matrix.
2. The high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 1, characterized in that The vector multiplication processing unit includes a preprocessing module and an adder tree module; The vector multiplication processing unit performs multiplication processing on the row elements of the any square matrix and the column elements of the ternary square matrix, and is specifically configured to: Calculate the element product of the row element and the column element through the preprocessing module; Calculate the product sum of the element products through the adder tree module and mark the product sum as a vector product.
3. A high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 2, characterized in that, Both the any square matrix and the ternary square matrix are of a preset dimension; the block matrix multiplication processing unit is further configured to: Calculate the number of vector products according to the preset dimension; Divide the any square matrix into multiple rows of elements, and divide the ternary square matrix into multiple columns of elements; Input the multiple rows of elements and the multiple columns of elements into a plurality of vector multiplication processing units corresponding to the number of vector products; wherein, one row of elements of the any square matrix and one column of elements of the ternary square matrix are input into each vector multiplication processing unit.
4. A high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 2, wherein, The preprocessing module performs the calculation of the element product of the row element and the column element, and is specifically configured to: Number the row elements according to the arrangement order of the row elements, and number the column elements according to the arrangement order of the column elements; Perform product calculation on the row element and the column element with the same number to obtain an element product.
5. A high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 4, characterized in that, The preprocessing module performs product calculation on the row element and the column element with the same number, and is specifically configured to: Obtain the target value corresponding to the column element; wherein, the column element of the ternary square matrix is a two-bit binary value; when the column element is 00, the target value corresponding to the column element is 0; when the column element is 01, the target value corresponding to the column element is 1; when the column element is 11, the target value corresponding to the column element is -1; Calculate the product of the row element and the target value and mark the product as the element product of the row element and the column element.
6. The high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 4, wherein The preprocessing module includes a preset number of bit calculation units; the preset number is the number of bits of the row element; The bit calculation unit includes an AND gate, a NOT gate, and a multiplexer; the output terminal of the AND gate is connected to the input terminal of the multiplexer, and the output terminal of the NOT gate is connected to the input terminal of the multiplexer; The preprocessing unit performs a multiplication calculation on the row elements and column elements with the same number, and is specifically configured to: Input a single-bit value of the row element into the AND gate and the NOT gate respectively, and input the least significant bit value of the column element into the AND gate; the column element is a two-bit binary value; Input the first processing result of the NOT gate and the second processing result of the AND gate into the multiplexer; Detect the most significant bit value of the column element; If the most significant bit value is 1, output the first processing result through the multiplexer and mark it as the bit product result corresponding to the single-bit value; if the most significant bit value is 0, output the second processing result through the multiplexer and mark it as the bit product result corresponding to the single-bit value; Combine the bit product results corresponding to the values of all bits of the row element to obtain the element product of the row element and the column element.
7. The high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 2, characterized in that The adder tree module calculates the sum of products of the element products, and is specifically configured to: Input multiple element products of multiple row elements and multiple column elements into the adder tree module; The adder tree module performs a compression process on the multiple element products and the sum of products of the previous clock cycle to generate an accumulated sum output for the current clock cycle; the sum of products of the previous clock cycle is the vector product obtained when the previous arbitrary square matrix is multiplied by the previous ternary square matrix; Mark the accumulated sum output of the current clock cycle as the sum of products of the element products.
8. A high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 7, characterized in that, The adder tree module includes a Wallace tree and a carry-ripple adder; the Wallace tree includes a plurality of carry-save adders, and the plurality of carry-save adders are used to compress multiple values into two values; The adder tree module performs a compression process on the multiple element products and the sum of products of the previous clock cycle, and is also configured to: Obtain the sum of products of the previous clock cycle; Process the multiple element products and the sum of products of the previous clock cycle through the Wallace tree, and fill the least significant bit of the carry output value of the carry-save adder with the most significant bit value of the column element to obtain two outputs of the Wallace tree.
9. A high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to claim 8, characterized in that, After the adder tree module obtains two outputs of the Wallace tree, it is also configured to: Use the most significant bit value of the column element as the carry input of the carry-ripple adder, and perform a summation process on the two outputs of the Wallace tree through the carry-ripple adder to generate an accumulated sum output for the current clock cycle.
10. A high-performance method for left and right multiplication of a ternary matrix and an arbitrary matrix, characterized in that, The multiplication processing method is applied to a high-performance device for left and right multiplication of a ternary matrix and an arbitrary matrix according to any one of claims 1-9; the method includes: Perform multiplication and accumulation processing on the row elements of an arbitrary square matrix and the column elements of a ternary square matrix to generate a vector product; the arbitrary square matrix is a block of an arbitrary matrix, and the ternary square matrix is a block of a ternary matrix; Generate the product matrix of the arbitrary square matrix and the ternary square matrix according to the obtained products of multiple vectors; Obtain the product result matrix of the arbitrary matrix and the ternary matrix according to the multiple product matrices of multiple arbitrary square matrices and multiple ternary square matrices.