Multiply-accumulate arithmetic unit and matrix multiplier comprising same
By designing a matrix multiplier with quantized weight matrix and using scaling factors and binary values, the problem of low deep learning computing efficiency in small semiconductor devices is solved, improving computing speed and circuit reliability, and reducing power consumption.
Patent Information
- Application Number
- CN202411480210.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2024-10-23
- Publication Date
- 2025-07-25
AI Technical Summary
The matrix multiplication operation required to achieve deep learning in small semiconductor devices is inefficient, resulting in insufficient computing power and difficult to meet the needs of mobile or IoT devices.
Design a matrix multiplier, including an input vector scaler, data type converter and multiplication accumulator array, and realize fixed-point and floating-point data conversion by quantizing the weight matrix and using scaling factors and binary values to perform operations, and improve calculation efficiency.
It reduces the area occupation of the matrix multiplier, improves the integrated density and reliability of the circuit, reduces power loss, and improves the calculation speed and yield rate.
Smart Images

Figure CN120372142A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority and benefit of Korean Patent Application No. 10 - 2024 - 0011775, filed with the Korean Intellectual Property Office on January 25, 2024, the entire content of which is incorporated herein by reference. Technical field
[0003] Various example embodiments relate to a multiply - accumulate (MAC) operator and / or a matrix multiplier including the operator. Background art
[0004] Deep learning in the field of artificial intelligence can identify patterns in complex data and achieve refined predictions. Generally, deep learning consists of (or includes) a training phase in which a neural network model is trained using training data and an inference phase in which new data is input into the trained neural network model to obtain an output. Most of the computing time of deep learning is spent on matrix multiplication.
[0005] Meanwhile, deep learning can generally be performed using general - purpose processors such as a graphics processing unit (GPU) and / or a neural processing unit (NPU). Recently, the trend of embedding deep - learning technology in mobile or Internet of Things (IoT) application devices is accelerating, and there is a need or desire for a technology that can implement the large amount of computing power required or used for deep learning in a small semiconductor. Summary of the invention
[0006] Various example embodiments may attempt to provide a multiply - accumulate (MAC) operator configured to be implemented in a small semiconductor device and / or a matrix multiplier including the operator.
[0007] Alternatively or additionally, various example embodiments may attempt to provide a multiply - accumulate (MAC) operator implemented as a single - piece hardware and configured to perform a multiplication operation on multiple input matrices, and / or perform a multiplication operation on an input matrix and a weight matrix, and / or a matrix multiplier including the operator.
[0008] A matrix multiplier according to various exemplary embodiments may include: an input vector scaler configured to generate a scaled input matrix based on a first input matrix and a plurality of scaling factors; a first data type converter configured to convert the data type of the scaled input matrix into fixed-point and generate a fixed-point input matrix; a multiply-accumulate arithmetic unit array configured to receive the fixed-point input matrix and a plurality of binary vectors, generate a fixed-point output matrix based on the fixed-point input matrix and the plurality of binary vectors, receive the first input matrix and a second input matrix, and generate a first output matrix based on the first input matrix and the second input matrix; and a second data type converter configured to convert the data type of the fixed-point output matrix into floating-point and generate a second output matrix.
[0009] Alternatively or additionally, a MAC arithmetic unit according to various exemplary embodiments may include: a first multiplexer configured to receive a first data of a first data type and mantissa data of a second data type different from the first data type, output the first data based on a first signal, and output the mantissa data of the second data type based on a second signal different from the first signal; a first AND arithmetic unit configured to receive a plurality of binary values and perform an AND operation of a first binary value among the plurality of binary values and the first data; and an adder configured to add the output data of the first AND arithmetic unit and output first output data of the first data type.
[0010] Alternatively or additionally, a matrix multiplier according to various exemplary embodiments may include: a data buffer configured to receive a first input vector and a second input vector from an external memory, and receive a plurality of scaling factors and a plurality of binary vectors as weight vectors from the external memory; an input parser configured to receive the plurality of scaling factors and the plurality of binary vectors, output the plurality of scaling factors and the plurality of binary vectors and a first signal of a first level, receive the second input vector, and output the second input vector and a first signal of a second level different from the first level; an input vector scaler configured to generate a first scaled input vector based on a multiplication operation of the first input vector and the plurality of scaling factors; a first data type converter configured to receive the first scaled input vector, extract a first exponent from exponents of each scaled input element among a first plurality of scaled input elements included in the first scaled input vector, and convert the data type of the first scaled input vector into fixed-point based on the first exponent; and a MAC arithmetic unit configured to accumulate the first scaled input vector based on binary values of the plurality of binary vectors and generate a fixed-point output matrix when receiving the first signal of the first level. Description of the Drawings
[0011] Figure 1It is a block diagram of a matrix multiplier according to some example embodiments.
[0012] Figure 2 Shows Figure 1 the weight matrix WM.
[0013] Figure 3 Shows floating-point data as a data format.
[0014] Figure 4 Shows a multiply-accumulate operator.
[0015] Figure 5 Shows Figure 4 the operation method of the mantissa operation unit of
[0016] Figure 6 Is a block diagram showing Figure 1 the configuration of the matrix multiplier of
[0017] Figure 7 It is a block diagram of an input vector scaler according to some example embodiments.
[0018] Figure 8 Shows Figure 7 the operation of the input vector scaling circuit of
[0019] Figure 9 It is a block diagram of a first data type converter according to some example embodiments.
[0020] Figure 10 Shows a method for extracting an exponent by Figure 9 the first data type converter of
[0021] Figure 11 Shows an operation method of converting the data type of a received input vector into fixed-point based on the exponent extracted by Figure 10 the first data type converter of
[0022] Figure 12 It is a block diagram of an array of multiply-accumulate operators according to some example embodiments.
[0023] Figure 13 Shows an operation method of a multiply-accumulate operator according to some example embodiments.
[0024] Figure 14 Shows an operation method of a multiply-accumulate operator according to some example embodiments.
[0025] Figure 15 Shows an operation method of a multiply-accumulate operator according to some example embodiments.
[0026] Figure 16It is a block diagram of a mantissa arithmetic unit of a multiply-accumulate arithmetic unit according to some example embodiments.
[0027] Figure 17 It is a block diagram of a multiply-accumulate arithmetic unit according to some example embodiments.
[0028] Figure 18 It shows the operation of a second data type converter according to some example embodiments.
[0029] Figure 19 It is a block diagram showing a neural processing system according to some example embodiments.
[0030] Figure 20 It shows an example of multi-head attention according to some example embodiments.
[0031] Figure 21 It shows Figure 20 an example of scaled dot-product attention. Detailed Description
[0032] In the following detailed description, only certain example embodiments are shown and described by way of illustration. As those of ordinary skill in the art can recognize, the described example embodiments can be modified in various different ways without departing from the spirit or scope of the inventive concept.
[0033] Accordingly, the drawings and the description are shown and described as illustrative in nature, and not restrictive. Throughout the specification, the same reference numerals designate the same elements. In the flowcharts described with reference to the drawings in this specification, the order of operations can be changed, several operations can be combined, certain operations can be divided, and specific operations can be not executed.
[0034] In the description, expressions described in the singular form in this specification can be construed as singular or plural unless an explicit expression such as "one" or "single" is used. Although terms such as "first", "second", etc. are used to explain various constituent elements, the constituent elements are not limited to such terms. These terms are only used to distinguish one constituent element from another.
[0035] Figure 1 It is a block diagram of a matrix multiplier according to some example embodiments. Referring to Figure 1 , the matrix multiplier MATMUL100 can receive a plurality of input matrices XM and YM. Each of the input matrices XM and YM can include a plurality of input vectors. Each input vector among the plurality of input vectors can include a plurality of input elements. For example, the first input matrix XM among the plurality of input matrices XM and YM can be represented by Equation 1 below.
[0036] (Equation 1)
[0037]
[0038] Referring to Equation 1, XM can represent the first input matrix XM, to can respectively represent the first input vector to the (h)th input vector, and x 11 to x hn can represent different input elements. For example, x 11 to x 1n can represent the input elements included in the first input vector , and x h1 to x hn can represent the input elements included in the (h)th input vector . In various example embodiments, h can be greater than, less than, or equal to n.
[0039] The second input matrix YM among the multiple input matrices XM and YM can be represented by Equation 2 below.
[0040] (Equation 2)
[0041]
[0042] Referring to Equation 2, YM can represent the second input matrix YM, and the second input matrix YM can include the first input vector to the (m)th input vector and each input vector can include different input elements from y 11 to y nm . For example, the first input vector can include different input elements from y 11 to y n1 , and the (m)th input vector can include different input elements from y 1m to y nm . In various example embodiments, m can be less than, greater than, or equal to n.
[0043] In some example embodiments, each input element or entry included in the input matrices XM and YM can be floating-point data. For example, the data type of each input element included in the input matrices XM and YM can be FP16 (16-bit floating point), or can be FP32 (32-bit floating point). However, the example embodiments are not limited thereto. In some cases, the data types of the input elements XM and YM can be heterogeneous; the example embodiments are not limited thereto.
[0044] In some example embodiments, the matrix multiplier 100 may perform matrix multiplication on a first input matrix XM and a second input matrix YM. The matrix multiplication may be based on the Strassen algorithm and / or the divide-and-conquer algorithm and / or various other matrix multiplication algorithms; the example embodiments are not limited thereto. The matrix multiplier 100 may perform a multiplication operation on the first input matrix XM and the second input matrix YM, and may generate a first output matrix ZM. For example, the matrix multiplier 100 may generate the first output matrix ZM by performing floating-point multiplication of the input elements of the first input matrix XM and the second input matrix YM, and then performing floating-point summation of the multiplication results. This will be referred to later Figure 6 and Figures 12 to 14 describe a method for the matrix multiplier 100 to perform multiplication of the first input matrix XM and the second input matrix YM according to some example embodiments.
[0045] The first output matrix ZM may include a plurality of output vectors. Each output vector among the plurality of output vectors may include a plurality of output elements. For example, the first output matrix ZM may be represented by Equation 3 below.
[0046] (Equation 3)
[0047]
[0048] In this case, ZM may represent the first output matrix ZM, and each of z 11 to z hm may represent a different output element. For example, z 11 to z 1m may represent the output elements included in the first output vector , and z h1 to z hm may represent the output elements included in the (h) output vector .
[0049] In some example embodiments, the matrix multiplier 100 may also receive a weight matrix WM. The weight matrix WM may include a plurality of weights. For example, the weight matrix WM may be represented by Equation 4 below.
[0050] (Equation 4)
[0051]
[0052] Referring to Equation 4, WM may represent the weight matrix WM, and each of w 11 to w nm may represent a different weight. For example, w ijIt may be the weight set in the (i)-th row and (j)-th column of the weight matrix WM. In some example embodiments, each weight included in the weight matrix WM may be floating-point data. For example, the data type of each weight may be FP16 and / or FP32. However, the example embodiments are not limited thereto.
[0053] In some example embodiments, the weights of the weight matrix WM may be approximated using a plurality of scaling factors (α) and binary values b. The matrix multiplier 100 may receive the plurality of scaling factors (α) and binary values b as the weights of the weight matrix WM. A detailed description of the plurality of scaling factors (α) and binary values b will be referred to later. Figure 2 A detailed description of the plurality of scaling factors (α) and binary values b.
[0054] In some example embodiments, the matrix multiplier 100 may perform matrix multiplication on the first input matrix XM and the weight matrix WM. For example, the matrix multiplier 100 may be implemented to calculate the output elements by first multiplying the input elements of the first input matrix XM by the scaling factor (α) and then accumulating the products of the operations based on the binary value b. A method for the matrix multiplier 100 to perform the multiplication of the first input matrix XM and the weight matrix WM according to some example embodiments will be referred to later. Figures 6 to 17 A description of the method for the matrix multiplier 100 to perform the multiplication of the first input matrix XM and the weight matrix WM according to some example embodiments.
[0055] In some example embodiments, the matrix multiplier 100 may multiply the first input matrix XM and the weight matrix WM to generate a second output matrix PM. The second output matrix PM may include a plurality of output vectors. Each output vector among the plurality of output vectors may include a plurality of output elements. For example, the second output matrix PM may be represented by Equation 5 below.
[0056] (Equation 5)
[0057]
[0058] In this case, PM may represent the second output matrix PM, and each of p 11 to p hm may represent different output elements. For example, p 11 to p 1m may represent the output elements included in the first output vector , and p h1 to p hm may represent the output elements included in the (h)-th output vector .
[0059] The matrix multiplier 100 according to some example embodiments may perform a multiplication operation on a plurality of input matrices XM and YM and output a first output matrix ZM. The matrix multiplier 100 according to some example embodiments may perform a multiplication operation on an input matrix XM and a weight matrix WM and output a second output matrix PM. The matrix multiplier 100 according to some example embodiments may be implemented as a single piece of hardware and may perform a multiplication operation on a plurality of input matrices and a multiplication operation on an input matrix and a weight matrix. Since a neural processing system for performing deep learning includes a matrix multiplier according to various example embodiments, the area occupied by the matrix multiplier in the neural processing system can be reduced.
[0060] When the area of the matrix multiplier is reduced, the integration density and / or power loss of a circuit implementing the matrix multiplier, for example, can be improved. Alternatively or additionally, the yield of a circuit implementing the matrix multiplier according to various example embodiments can be increased. Alternatively or additionally, the reliability of a circuit implementing the matrix multiplier according to various example embodiments can be increased.
[0061] Figure 2 The weight matrix WM of Figure 1 is shown. As described with reference to Figure 1 , the matrix multiplier 100 may receive a plurality of scaling factors (α) and binary values b as weights of the weight matrix WM.
[0062] Through quantization of the weight matrix WM, each of the plurality of weights may be converted into a plurality of scaling factors (α) and binary values b. In various example embodiments, in order to increase the calculation speed of the first input matrix XM and the weight matrix WM, a quantization operation may be performed on the weight matrix WM. Through the quantization operation of the weight matrix WM, the weights included in the weight matrix WM may be approximated to a plurality of quantization levels QL. The number of the plurality of quantization levels QL may be determined by a predetermined quantization resolution. For example, each of the plurality of weights may be approximated to 2 R (where R is the quantization resolution) quantization levels QL through the quantization operation. In this case, each of the 2 R quantization levels QL may be determined based on a combination of R scaling factors (α) and R binary values b.
[0063] In some example embodiments, the weight matrix WM may include a plurality of row vectors. Referring to Figure 2 , the weight matrix WM may include a first row vector to an (n)th row vector In some example embodiments, the weight matrix WM may be quantized row by row. Specifically, α k_rj may represent the (j)th row vector for the weight matrix WM The (k)-th scaling factor. In this case, the scaling factor for the weights included in the first row of the weight matrix WM may be different from the scaling factor for the weights included in the second row of the weight matrix WM. In some example embodiments, each of the plurality of scaling factors (α) may have the same data type as the weights of the weight matrix WM. For example, the data type of each of the plurality of scaling factors (α) may be FP16 and / or FP32. However, the example embodiments are not limited thereto.
[0064] In some example embodiments, may represent the binary vector corresponding to the (k)-th scaling factor α of the (j)-th row vector k_rj corresponding thereto. is the sign value corresponding to the (k)-th scaling factor α of the (j)-th row vector k_rj corresponding thereto, and may include a plurality of binary values b.
[0065] The binary vector may be implemented as a row vector having the same dimension as the number of columns of the weight matrix WM. For example, may be represented by Equation 6 below.
[0066] (Equation 6)
[0067]
[0068] b 1_k_rj to b m_k_rj may each represent a different binary value b, for example, a value having an entry selected from {"0", "1"}. Specifically, b 1_k_rj to b m_k_rj may each be the binary value b for the weights set in different columns of the weight matrix WM. b 1_k_rj to b n_k_rj may each be "0" or "1".
[0069] As described above, each weight of the weight matrix WM may be approximated based on a plurality of scaling factors (α) and a plurality of binary values b. Here, quantization of each row of the weight matrix WM has been described, but is not limited thereto, and quantization may be performed on each column of the weight matrix WM.
[0070] In some example embodiments, the matrix multiplier 100 may receive a plurality of scaling factors (α) and binary values b as weights of the weight matrix WM. In some example embodiments, the matrix multiplier 100 may perform a multiplication operation on the first input matrix XM and the plurality of scaling factors (α) and binary values b, and may generate a second output matrix PM. For example, the matrix multiplier 100 may be implemented to output an element by first multiplying an element of the first input matrix XM by the scaling factor (α), and then accumulating the product of the operation based on the binary value b. The detailed operation method of the matrix multiplier 100 will be described later with reference to Figures 6 to 17 Describe the detailed operation method of this matrix multiplier 100.
[0071] Figure 3 and Figure 4 Illustrates the operation method of the multiply-accumulate unit. Figure 3 Illustrates floating-point data as a data format, and Figure 4 Illustrates the multiply-accumulate unit.
[0072] As described above, each input element of the input matrices XM and YM and the weights of the weight matrix WM may be floating-point data.
[0073] Reference Figure 3 , floating-point data may appear in various forms depending on the precision, but it can also be implemented using the IEEE 754 standard. The IEEE 754 standard represents real numbers by dividing them into a sign, an exponent, and a mantissa. For example, floating-point data can be represented by Equation 7 below.
[0074] (Equation 7)
[0075] x = -1 s × mantissa × 2 e
[0076] In single precision, floating-point data uses a total of 32 bits (FP32). Specifically, the sign (e.g., the most significant bit) is a single bit representing the sign of the data, and this bit can be "0" when the sign is positive and "1" when the sign is negative. The exponent can be 8 bits and represents the exponent, while the mantissa can be 23 bits (the least significant) and represents the mantissa or significant digits.
[0077] For example, representing -314.625 in IEEE 754 floating-point format is as follows:
[0078] First, since the sign is negative, s becomes "1".
[0079] Regarding the mantissa, if the absolute value of the number 314.625 is represented in binary, it becomes 100111010.101 (2)At this time, if the decimal point is moved to the left so that only 1 remains on the left side of the decimal point, it can be represented as follows, and this is called the normalized representation method.
[0080] 100111010.101₂ = 1.00111010101₂ × 2 8
[0081] Here, the mantissa is the part to the right of the decimal point, that is, 00111010101.
[0082] Regarding the exponent, in the normalized representation of -314.625, the exponent is 8, and adding the bias value of 127 to the exponent 8 becomes 135. At this time, when 135 is converted to binary, it becomes 10000111 (2) , and this value becomes the exponent.
[0083] Although single precision is described as an example here, in half precision or double precision, the number of bits of the exponent and mantissa and the bias value can be different. In the following, it will be described under the assumption that the data type of the input elements of the input matrices XM and YM and the scaling factor (α) of the weight matrix WM are FP32.
[0084] Figure 4 Shows a multiply-accumulate arithmetic unit.
[0085] The multiply-accumulate arithmetic unit 400 can receive floating-point data as input data and perform a multiply-accumulate operation on the input data. In the following, the operation method of the multiply-accumulate arithmetic unit 400 will be described in detail.
[0086] Reference Figure 4 , the multiply-accumulate arithmetic unit 400 can receive the first input data IN1[31:0] and the second input data IN2[31:0], and can perform a multiply-accumulate operation on the input data. For example, when the first input data IN1[31:0] and the second input data IN2[31:0] are 32-bit floating-point data, the first input data IN1[31:0] and the second input data IN2[31:0] can each include 1-bit sign data S_IN1
[31] and S_IN2
[31] , 8-bit exponent data E_IN1[30:23] and E_IN2[30:23], and 23-bit mantissa data M_IN1[22:0] and M_IN2[22:0].
[0087] The multiply-accumulate arithmetic unit 400 can include a multiplication part 401 and an accumulation part 402. The multiplication part 401 can include a sign operation unit 410, an exponent operation unit 420, and a mantissa operation unit 430, and the accumulation part 402 can include an adder 450 and an accumulator 460.
[0088] Looking at the multiplication part 401, the sign operation unit 410 may include an exclusive - or (XOR) operation unit 411. The sign operation unit 410 may receive the first sign data S_IN1
[31] of the first input data IN1 and the second sign data S_IN2
[31] of the second input data IN2. When both the first sign data S_IN1
[31] and the second sign data S_IN2
[31] have the value "0" representing a positive number or the value "1" representing a negative number, the XOR operation unit 411 of the sign operation unit 410 may output "0" representing a positive number (e.g., because a positive number multiplied by a positive number is positive, and a negative number multiplied by a negative number is positive). On the other hand, when one of the first sign data S_IN1
[31] and the second sign data S_IN2
[31] has the value "0" representing a positive number and the other has the value "1" representing a negative number, the XOR operation unit 411 of the sign operation unit 410 may output "1" representing a negative number (e.g., because a positive number multiplied by a negative number is negative, and a negative number multiplied by a positive number is negative). The sign operation unit 410 may output the data generated as a result of the XOR operation as 1 - bit sign data OUT
[31] of the output data OUT.
[0089] The exponent operation unit 420 may include a first exponent adder 421 and a second exponent adder 423. The first exponent adder 421 may receive the first exponent data E_IN1[30:23] of the first input data IN1 and the second exponent data E_IN2[30:23] of the second input data IN2. The first exponent adder 421 may perform a first addition operation on the first exponent data E_IN1[30:23] and the second exponent data E_IN2[30:23]. The first exponent adder 421 may output the first addition result data generated as a result of the first addition operation. The first exponent data E_IN1[30:23] and the second exponent data E_IN2[30:23] are each added to an exponent bias value (e.g., "127"). Since the exponent bias value is doubled through the first addition operation in the first exponent adder 421, the exemplary embodiment subtracts the exponent bias value from the first addition result data. Accordingly, the second exponent adder 423 may receive the first addition result data output from the first exponent adder 421 and perform a second addition operation of subtracting the exponent bias value (e.g., "127") from the first addition result data. The second exponent adder 423 may output the data generated as a result of the second addition operation as 8 - bit exponent data E_OUT[7:0].
[0090] The mantissa operation unit MTS_MUL 430 can perform a multiplication operation on mantissa data. The mantissa operation unit 430 can receive the first mantissa data M_IN1[22:0] of the first input data IN1 and the second mantissa data M_IN2[22:0] of the second input data IN2. The mantissa operation unit 430 can perform a multiplication operation on the first mantissa data M_IN1[22:0] and the second mantissa data M_IN2[22:0]. The mantissa operation unit 430 can output the data generated as a result of the multiplication operation as 48-bit mantissa data M_OUT[47:0]. The specific operation method of the mantissa operation unit 430 will be described later with reference to Figure 5 Describe the specific operation method of the mantissa operation unit 430.
[0091] The normalizer 440 can normalize the exponent data E_OUT[7:0] output from the exponent operation unit 420 and the mantissa data M_OUT[47:0] output from the mantissa operation unit 430. Normalization can refer to shifting the calculated data so that the most significant bit becomes 1. The normalized data OUT[30:0] can be output to the accumulation part 402 together with the sign data OUT
[31] output from the sign processing circuit 410 as the multiplication result value OUT.
[0092] Looking at the accumulation part 402, the adder 450 can perform an accumulation operation by adding the data received from the multiplication part 401 and the data received from the accumulator 460.
[0093] Figure 5 Shows Figure 4 The operation method of the mantissa operation unit. Generally, binary multiplication can be performed by shift and addition operations. The mantissa operation unit 430 that performs binary multiplication can be implemented as a shifter 431, an AND operation unit 432, and an adder 433.
[0094] Refer to together Figure 4 And Figure 5 And, the mantissa operation unit 430 can receive the first mantissa data M_IN1[22:0] as the multiplicand and receive the second mantissa data M_IN2[22:0] as the multiplier. Based on each bit M_IN2[0], M_IN2[1], ……, M_IN2
[22] , the data obtained from the first mantissa data M_IN1[22:0] and the second mantissa data M_IN2[22:0] can be input to the adder 433 after AND operation.
[0095] The adder 433 can output the mantissa data M_OUT[47:0].
[0096] In some example embodiments, the least significant bit M_IN2[0] of the first mantissa data M_IN1[22:0] and the second mantissa data M_IN2[22:0] may be input to the first AND operation unit 432_1. The first AND operation unit 432_1 may output the first mantissa data M_IN1[22:0] as it is, or output data with all bits being "0", based on the least significant bit M_IN2[0] of the second mantissa data M_IN2[22:0]. For example, if the least significant bit M_IN2[0] of the second mantissa data M_IN2[22:0] is "1", the first AND operation unit 432_1 may output the first mantissa data M_IN1[22:0] as it is. If the least significant bit M_IN2[0] of the second mantissa data M_IN2[22:0] is "0", the first AND operation unit 432_1 may output data with all bits being "0".
[0097] In some example embodiments, the shifter 431 may shift the first mantissa data M_IN1[22:0] left by 1 bit. The first mantissa data M_IN1[22:0] shifted left by the shifter 421 may be input to the AND operation units 432_2, ……, 432_3 together with each bit M_IN2[1], ……, M_IN2
[22] of the second mantissa data M_IN2[22:0].
[0098] In various example embodiments, the data output by the AND operation unit 432 may be input to the adder 433. The adder 433 may perform an addition operation on the input data sequentially or simultaneously at once. The adder 433 may output 48-bit mantissa data M_OUT[47:0] as the multiplication result of the first mantissa data M_IN1[22:0] and the second mantissa data M_IN2[22:0]. Here, the mantissa operation unit 430 is shown to include one adder, but is not limited thereto, and the mantissa operation unit 430 may include a plurality of adders connected to each AND operation unit 433.
[0099] Figure 6 is a block diagram showing Figure 1 the configuration of the matrix multiplier. Specifically, Figure 1 the matrix multiplier 100 of
[0100] In some example embodiments, the matrix multiplier 600 may include a first input matrix buffer (input matrix buffer #1 610), a second input matrix buffer (input matrix buffer #2 620), a weight buffer 630, an input parser 640, an input vector scaler 650, a first data type converter (data type converter #1 660), a multiply-accumulate operation unit array (MAC_ARRAY 670), and a second data type converter (data type converter #2 680).
[0101] In various example embodiments, the first input matrix buffer 610 and the second input matrix buffer 620 may store input matrices provided from the outside. For example, the first input matrix buffer 610 and the second input matrix buffer 620 may receive a first input matrix XM and a second input matrix YM from an external memory, respectively, and may provide the first input matrix XM and the second input matrix YM to the input vector scaler 650 and the multiply-accumulate operation unit array 670.
[0102] In some example embodiments, the weight buffer 630 may store a plurality of scaling factors (α) and binary values b provided from the outside. The weight buffer 630 may store the plurality of scaling factors (α) and the binary values b as weights of a weight matrix WM. For example, the weight buffer 630 may receive the plurality of scaling factors (α) and the binary values b from the outside, and may provide the plurality of scaling factors (α) and the binary values b to the input vector scaler 650 and the multiply-accumulate operation unit array 670.
[0103] In some example embodiments, the input parser 640 may receive the second input matrix YM from the second input matrix buffer 620, or may receive the plurality of scaling factors (α) and the binary values b from the weight buffer 630. The input parser 640 may parse the second input matrix YM or the plurality of scaling factors (α) and the binary values b, and may provide the plurality of scaling factors (α) and the binary values b to the input vector scaler 650 or the multiply-accumulate operation unit array 670.
[0104] In some example embodiments, the second input matrix buffer 620 and the weight buffer 630 may be implemented as one data buffer. Hereinafter, it is assumed that the input parser 640 receives the second input matrix YM from the second input matrix buffer 620, and receives the plurality of scaling factors (α) and the binary values b from the weight buffer 630, but the example embodiments are not limited thereto. The input parser 640 may receive the second input matrix YM, the plurality of scaling factors (α), and the binary values b from the data buffer implemented as one.
[0105] In some example embodiments, when the input parser 640 receives the second input matrix YM from the second input matrix buffer 620, the input parser 640 may provide the second input matrix YM to the multiply-accumulate array 670 such that the multiply-accumulate units in the multiply-accumulate array 670 perform matrix multiplication on the first input matrix XM and the second input matrix YM.
[0106] In some example embodiments, when the input parser 640 receives a plurality of scaling factors (α) and binary values b from the weight buffer 630, the input parser 640 may provide the plurality of scaling factors (α) to the input vector scaler 650 and provide the plurality of binary values b to the multiply-accumulate array 670 such that a multiplication operation is performed on the first input matrix (XM), the plurality of scaling factors (α), and the binary values b. The input vector scaler 650 may perform a multiplication operation on the plurality of scaling factors (α) and the first input matrix XM and output a scaled input matrix XM'.
[0107] In some example embodiments, the multiply-accumulate array 670 may include a plurality of multiply-accumulate units. The multiply-accumulate array 670 may receive from the first data type converter 660 a fixed-point input matrix XM′_ in which the data type of the scaled input matrix XM' is converted to fixed-point. fxp In some example embodiments, the input parser 640 may output a control signal SEL of a first level (e.g., high level) to the multiply-accumulate array 670 such that the multiply-accumulate units in the multiply-accumulate array 670 perform a multiplication operation on the fixed-point input matrix XM′_ fxp and the binary values b.
[0108] The multiply-accumulate array 670 may perform a multiplication operation on the fixed-point input matrix XM′_ fxp and the binary values b by accumulating the elements of the fixed-point input matrix XM′_ fxp based on the binary values b and may output a fixed-point output matrix PM_ fxp .
[0109] In some example embodiments, the input parser 640 may output a control signal SEL of a second level (e.g., low level) to the multiply-accumulate array 670 such that the multiply-accumulate units in the multiply-accumulate array 670 perform matrix multiplication on the first input matrix XM and the second input matrix YM. The multiply-accumulate array 670 may perform a multiply-accumulate operation on the input elements of the first input matrix XM and the input elements of the second input matrix YM and output a first output matrix ZM.
[0110] In some example embodiments, the input parser 640 may output a control signal SEL to the second data type converter 680 such that the second data type converter 680 converts the data format of the output elements of the multiply-accumulate arithmetic unit array 670. The second data type converter 680 may convert the fixed-point output matrix PM_ fxp into a second output matrix PM in floating-point data format.
[0111] In some example embodiments, the input vector scaler 650 may receive a first input matrix XM. For example, the input vector scaler 650 may receive a plurality of input vectors including a plurality of input elements (e.g., to ). The input vector scaler 650 may scale the first input matrix XM based on a plurality of scaling factors (α). For example, the input vector scaler 650 may generate a plurality of scaled input vectors based on the plurality of input vectors. In this case, the plurality of scaled input vectors may respectively correspond to the plurality of input vectors to Hereinafter, the scaled input vectors corresponding to to will be respectively referred to as to The specific configuration and operation method of the input vector scaler 650 will be described later with reference to Figure 7 and Figure 8
[0112] In some example embodiments, the first data type converter 660 may extract an exponent EXP from each of the scaled input vectors . For example, the first data type converter 660 may extract a first exponent EXP1 from the first scaled input vector and extract a second exponent EXP2 from the second scaled input vector . The first data type converter 660 may provide the extracted exponent EXP to the second data type converter 680.
[0113] In some example embodiments, the first data type converter 660 may convert the data type of the plurality of scaled input vectors to fixed-point. For example, the first data type converter 660 may receive the plurality of scaled input vectors and output a plurality of fixed-point input vectors The first data type converter 660 may convert the data type of the scaled input elements included in the plurality of scaled input vectors to fixed-point based on the extracted exponent EXP. The specific content will be described later with reference to Figures 9 to 11Describe the specific configuration and operation method of the first data type converter 660.
[0114] In some example embodiments, the multiply-accumulate (MAC) array 670 may receive a first input matrix XM from the first input matrix buffer 610 and a second input matrix YM from the input parser 640. The MAC array 670 may also receive a control signal SEL from the input parser 640. The MAC array 670 may perform matrix multiplication on the first input matrix XM and the second input matrix YM based on the control signal SEL at a second level and output a first output matrix ZM. In various example embodiments, the MAC array 670 may perform multiply-accumulate operations on the input elements of the first input matrix XM and the input elements of the second input matrix YM and generate the output elements of the first output matrix ZM.
[0115] In some example embodiments, the MAC array 670 may receive a plurality of fixed-point input vectors from the first data type converter 660 and a binary vector including a plurality of binary values b from the input parser 640 The MAC array 670 may receive a control signal SEL at a first level from the input parser 640. The MAC array 670 may perform a multiplication operation on the fixed-point input vectors by accumulating the elements of the fixed-point input vectors and the binary values b and generate a fixed-point output vector
[0116] Reference will be made later to Figures 12 to 16 Describe the detailed configuration and operation method of the MAC array 670.
[0117] In some example embodiments, the second data type converter 680 may receive a plurality of exponents EXP from the first data type converter 660. The second data type converter 680 may receive the fixed-point output vector from the MAC array 670 The second data type converter 680 may receive a plurality of control signals SEL at a first level and convert the data type of the fixed-point output vector based on the exponents EXP to floating point. For example, the second data type converter 680 may generate a second output vector including floating-point output elements In some example embodiments, the second data type converter 680 may receive the first output matrix ZM from the MAC array 670. The second data type converter 680 may receive a control signal SEL at a second level and output the first output matrix ZM as the output data of the second data type converter 680. Reference will be made later toFigure 18 Describe the specific configuration and operation method of the second data type converter 680.
[0118] Figure 7 is a block diagram of an input vector scaler according to some example embodiments. The input vector scaler 650 according to some example embodiments may perform a multiplication operation on a first input matrix XM and a plurality of scaling factors (α). In some example embodiments, the input vector scaler 650 may include a first input vector scaling circuit 651 to an (h)th input vector scaling circuit 65h.
[0119] In some example embodiments, each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h may receive a different input vector. For example, the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h may receive a first input vector to an (h)th input vector (i.e., to ).
[0120] In some example embodiments, each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h may sequentially receive a plurality of input elements. For example, the first input vector scaling circuit 651 may sequentially receive x 11 to x 1n , and the (h)th input vector scaling circuit 65h may sequentially receive x h1 to x hn .
[0121] In some example embodiments, each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h may receive a plurality of scaling factors (α) from the input parser 640. The input parser 640 may parse a plurality of scaling factors (α) received from the outside and sequentially provide the plurality of scaling factors (α) to the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h.
[0122] In some example embodiments, the plurality of scaling factors (α) received by each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h from the input parser 640 may be the same. For example, the plurality of scaling factors (α) received by the first input vector scaling circuit 651 may be the same as the plurality of scaling factors (α) received by the second input vector scaling circuit 652.
[0123] In some example embodiments, the order in which each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h receives a plurality of scaling factors (α) may be the same. For example, the scaling factor (α) initially received by the first input vector scaling circuit 651 may be the same as the scaling factor (α) initially received by the second input vector scaling circuit 652. Similarly, the scaling factor (α) received by the first input vector scaling circuit 651 the second time may be the same as the scaling factor (α) received by the second input vector scaling circuit 652 the second time.
[0124] In some (first) example embodiments, each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h may perform a multiplication operation on the input elements and the plurality of scaling factors (α) based on the order of receiving the input elements and the plurality of scaling factors (α). Since the operation method of each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h is the same as or similar to Figure 4 the operation method of the multiply-accumulate operator described in, the detailed description of the operation method of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h is omitted here.
[0125] In some example embodiments, each of the first input vector scaling circuit 651 to the (h)th input vector scaling circuit 65h may perform a multiplication operation on the input elements and the plurality of scaling factors (α), and each may generate a first scaled input vector to the (h)th scaled input vector For example, the first input vector scaling circuit 651 may generate and the second input vector scaling circuit 652 may generate
[0126] Figure 8 shows Figure 7 the operation of the input vector scaling circuit. Hereinafter, for a more concise description, the operation of the first input vector scaling circuit 651 will be representatively described. However, the example embodiments are not limited thereto, and the second input vector scaling circuit 652 to the (h)th input vector scaling circuit 65h may also operate in the same or similar manner.
[0127] In some example embodiments, the first input vector scaling circuit 651 may receive a first input vector That is, the first input vector scaling circuit 651 may sequentially receive x 11 to x 1n .
[0128] In some example embodiments, the first input vector scaling circuit 651 may sequentially receive a plurality of scaling factors (α). For example, the first input vector scaling circuit 651 may receive the scaling factor α corresponding to the first row vector of the weight matrix WM 1_r1 to α R_r1 , and then may receive the scaling factor α corresponding to the second row vector 1_r2 to α R_r2 . In this way, the first input vector scaling circuit 651 may sequentially receive all the plurality of scaling factors α 1_r1 to α R_rn from the input parser 640.
[0129] In some example embodiments, the first input vector scaling circuit 651 may perform a multiplication operation on the input elements x 11 to x 1n and the plurality of scaling factors (α). In some example embodiments, the first input vector scaling circuit 651 may perform operations on the sign data, exponent data, and mantissa data of the input elements x 11 to x 1n respectively, and the plurality of scaling factors (α).
[0130] In some example embodiments, the first input vector scaling circuit 651 may output the scaled elements as a first scaled input vector based on the multiplication operation on the input elements x 11 to x 1n and the plurality of scaling factors (α). For example, the first input vector scaling circuit 651 may sequentially multiply the received input elements by the plurality of scaling factors (α) to calculate the elements of the plurality of scaled input vectors. The first input vector scaling circuit 651 may multiply x 11 by the scaling factor α corresponding to the first row vector of the weight matrix WM 1_r1 to α R_r1 to sequentially calculate the plurality of scaled input elements 651_1 corresponding to x 11 and x 11 ×α k_r1 .
[0131] Thereafter, the first input vector scaling circuit 651 may multiply x 12 by the scaling factor α corresponding to the second row vector of the weight matrix WM 1_r2 to α R_r2 to sequentially calculate the plurality of scaled input elements 651_2 and x 12 corresponding to x 12 ×αk_r2 In this way, the first input vector scaling circuit 651 can operate sequentially until multiple scaled input elements 651_3 corresponding to x 1n and x 1n ×α k_rn . Hereinafter, multiple scaled input elements 651_1 corresponding to x 11 can be represented by x'. 11 Similarly, multiple scaled input elements 651_2 corresponding to x 12 can be represented by x', and multiple scaled input elements 651_3 corresponding to x 12 can be represented by x'. 1n can be represented by x'. 1n .
[0132] In some example embodiments, the first input vector scaling circuit 651 can sequentially output multiple scaled input elements 651_1, 651_2,..., 651_3.
[0133] Figure 9 is a block diagram of a first data type converter according to some example embodiments. The first data type converter 660 according to some example embodiments can convert the data type of multiple scaled input vectors to a fixed point. In some example embodiments, the first data type converter 660 can include a first data type conversion circuit 661 to an (h)th data type conversion circuit 66h.
[0134] In some example embodiments, each of the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can receive a different input vector. For example, the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can respectively receive a first scaled input vector to an (h)th scaled input vector
[0135] In some example embodiments, each of the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can extract an exponent from multiple input elements included in the scaled input vector. For example, the first data type conversion circuit 661 can extract a first exponent EXP1 from the input elements (i.e., x' 11 to x' 1n ) included in the first scaled input vector, and the second data type conversion circuit 662 can extract an exponent from the input elements (i.e., x' 21 to x' 2n)Extract the second exponent EXP2. In this way, the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can respectively extract the first exponent EXP1 to the (h)th exponent EXPh. Each of the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can provide the extracted exponent to the second data type converter 680.
[0136] In some example embodiments, each of the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can convert the data type of the scaled input vector to fixed-point based on the extracted exponents EXP1 to EXPh. The first data type conversion circuit 661 to the (h)th data type conversion circuit 66h can respectively output the first fixed-point input vector to the (h)th fixed-point input vector For example, the first data type conversion circuit 661 can convert each of the input elements x′ included in the first scaled input vector 11 to x′ 1n to the fixed-point format based on the first exponent EXP1. The detailed operations of the first data type conversion circuit 661 to the (h)th data type conversion circuit 66h will be described below with reference to Figure 10 and Figure 11 .
[0137] Figure 10 shows a method for extracting an exponent by the first data type converter of Figure 9 . Hereinafter, for more concise explanation, the operation of the first data type conversion circuit 661 extracting the exponent will be representatively described. However, the example embodiments are not limited thereto.
[0138] Referring to Figure 10 , the first data type conversion circuit 661 can receive a plurality of scaled input elements. For example, the first data type conversion circuit 661 can receive x′ 11 to x′ 1n .
[0139] In some example embodiments, the data type of each input element among the plurality of input elements can be floating-point. For example, x′ 11 to x′ 1n can each include sign data S, exponent data EXP, and mantissa data MTS.
[0140] In some example embodiments, the first data type conversion circuit 661 may determine the maximum value among the exponent EXP values of each of the multiple received input elements. In this case, the first data type conversion circuit 661 may determine the exponent of the determined input element as the first exponent EXP1. That is, the first data type conversion circuit 661 may extract the exponent x′ 11 to x′ 1n and find the maximum exponent among them.
[0141] For more concise description, some example embodiments in which the first data type conversion circuit 661 extracts the maximum value among the exponent values of multiple input elements are representatively described in Figure 10 , but the example embodiments are not limited thereto. For example, the first data type conversion circuit 661 may be implemented to extract the minimum value among the exponent values of multiple input elements.
[0142] Figure 11 FIG. shows an operation method of converting the data type of the received input vector into fixed-point based on the exponent extracted by the Figure 10 first data type converter. Hereinafter, for more concise description, the operation of the first data type conversion circuit 661 extracting the exponent will be representatively described. However, the example embodiments are not limited thereto.
[0143] Refer to Figure 11 , the first data type conversion circuit 661 may receive multiple input elements. For example, the first data type conversion circuit 661 may receive x′ 11 to x′ 1n .
[0144] In some example embodiments, the first data type conversion circuit 661 may extract the first exponent EXP1. The first data type conversion circuit 661 may convert the data type of each scaled input element into fixed-point based on the first exponent EXP1. For example, the first data type conversion circuit 661 may convert x′ 11 to x′ 1n into x′ 11_fxp to x′ 1n_fxp respectively. However, hereinafter, for more concise description, the operation of the first data type conversion circuit 661 that converts the input element "x′ 11 " into the fixed-point input element "x′ 11_fxp " will be representatively described.
[0145] In some example embodiments, the first data type conversion circuit 661 may shift the scaled input element "x′ 11the mantissa of “” (hereinafter referred to as “the first mantissa MTSa”) is shifted by sft1, the difference between the first exponent EXP1 and the exponent of x′ 11 between the exponents. For example, when the difference between the first exponent EXP1 and the exponent value of x′ 11 EXP is “4”, the first data type conversion circuit 661 can insert “4” “0” bits into the most significant bit (MSB) position of the first mantissa MTSa.
[0146] In some example embodiments, the first data type conversion circuit 661 can determine the fixed-point input element “x′ 11_fxp ” mantissa (hereinafter referred to as “the second mantissa MTSb”) based on the shifted first mantissa MTSa. For example, the first data type conversion circuit 661 can truncate the lower-order bits of the shifted first mantissa MTSa according to the number of bits of the second mantissa MTSb. Alternatively, depending on the number of bits of the second exponent MTSb, the first data type conversion circuit 661 can determine the lower-order bits of the shifted first mantissa MTSa based on various types of rounding algorithms (such as “round to nearest even”). However, the example embodiments are not limited thereto.
[0147] In some example embodiments, the number of bits of the first mantissa MTSa can be “23 bits”. However, the example embodiments are not limited thereto.
[0148] In some example embodiments, the number of bits of the second mantissa MTSb can be “7 bits”. However, the example embodiments are not limited thereto.
[0149] In some example embodiments, each fixed-point input element among the multiple fixed-point input elements can be an integer. For example, the data type of each fixed-point input element among the multiple fixed-point input elements can be INT8 (8-bit integer). However, the example embodiments are not limited thereto.
[0150] Figure 12 is a block diagram of a multiply-accumulate arithmetic unit array according to some example embodiments. In some example embodiments, the multiply-accumulate arithmetic unit array 670 can include a plurality of multiply-accumulate arithmetic units MAC arranged in row and column directions. Hereinafter, for a more concise description, it is assumed that the plurality of multiply-accumulate arithmetic units MAC are arranged along (h) rows and (m) columns.
[0151] In some example embodiments, the multiply-accumulate (MAC) array 670 may include a first MAC row MR1 to an (h)-th MAC row MRh. Each of the first MAC row MR1 to the (h)-th MAC row MRh may include a plurality of multiply-accumulate units (MACs). For example, the first MAC row MR1 may include a plurality of multiply-accumulate units 671 to 67m. In some example embodiments, each of the first MAC processor MR1 to the (h)-th MAC processor MRh may receive an input vector of the first input matrix XM to and a fixed-point input vector of the fixed-point input matrix XM′_ fxp
[0152] In some example embodiments, different MAC rows may receive different input vectors to and different fixed-point input vectors In some example embodiments, the MACs included in the same MAC row may receive the same input vector to and the same fixed-point input vector For example, each of the multiply-accumulate units 671 to 67m included in the first MAC row MR1 may receive a first input vector and a first fixed-point input vector
[0153] In some example embodiments, the multiply-accumulate (MAC) array 670 may include a first MAC column MC1 to an (m)-th MAC column MCm. Each of the first MAC column MC1 to the (m)-th MAC column MCm may include a plurality of multiply-accumulate units (MACs). In some example embodiments, each of the first MAC column MC1 to the (m)-th MAC column MCm may receive an input vector of the second input matrix YM to and a binary value b of the binary vector In some example embodiments, different MAC columns may receive different input vectors to and different binary values b1, b2,..., b of the binary vector m . In some example embodiments, the MACs included in the same MAC column may receive the same input vector to and the binary vector with the same binary values b1, b2, ……, b m . For example, the multiply-accumulate (MAC) units included in the first column of MAC units MC1 can receive the first input vector and the binary vector with the first binary value b1.
[0154] In some example embodiments, each MAC unit can perform a multiply-accumulate operation on the input elements. For example, the MAC unit 671 can perform a multiply-accumulate operation on the elements of the first input vector of the first input matrix XM and the elements of the first input vector of the second input matrix YM, and output the first element z 11 of the first output matrix ZM.
[0155] Reference will be made to Figure 13 and Figure 14 to describe the specific operation methods of these MAC units.
[0156] In some example embodiments, each MAC unit can perform a multiplication operation on the fixed-point input elements and the binary values. For example, the MAC unit 671 can perform a multiplication operation on the fixed-point input elements and the binary values based on the first binary value b1 of the binary vector by accumulating the input elements of the fixed-point input vector , and output the first fixed-point output element p 11_fxp . Reference will be made to Figures 15 to 17 to describe the specific operation methods of these MAC units.
[0157] Figure 13 and Figure 14 illustrate the operation methods of the MAC units according to some example embodiments. Specifically, the MAC units can perform a multiply-accumulate operation on the first input matrix XM and the second input matrix YM, and output the first output matrix ZM.
[0158] In some example embodiments, the MAC units can perform a multiply-accumulate operation on the input elements included in the first row 1311 of the first input matrix XM and the input elements included in the first column 1412 of the second input matrix YM to calculate the first output element z 11 included in the first output matrix ZM.
[0159] In addition, the multiply-accumulate operator may perform a multiply-accumulate operation on the input elements included in the first row 1311 of the first input matrix XM and the input elements included in the second column 1422 of the second input matrix YM to calculate the second output element z included in the first output matrix ZM 12 。
[0160] Referring together Figure 14 ,in some example embodiments, each of the multiple multiply-accumulate operators 671, 672, ……, 67m included in the first multiply-accumulate operator row MR1 may receive a first input vector of the first input matrix XM For example, each of the multiple multiply-accumulate operators 671, 672, ... 67m may sequentially receive the elements x of the first input vector 11 to x 1n 。
[0161] In some example embodiments, each of the multiple multiply-accumulate operators 671, 672, ……, 67m included in the first multiply-accumulate operator row MR1 may receive a different input vector of the second input matrix YM to For example, the first multiply-accumulate operator 671 may sequentially receive multiple input elements y included in the first input vector of the second input matrix YM 11 to y n1 ,and the second multiply-accumulate operator 672 may sequentially receive multiple input elements y included in the second input vector of the second input matrix YM 12 to y n2 。
[0162] In some example embodiments, each multiply-accumulate operator may perform a multiply-accumulate operation on these input elements based on the order of receiving the input elements of the first input matrix XM and the input elements of the second input matrix YM. For example, the first multiply-accumulate operator 671 may calculate the first output element z according to Equation 8 below 11 。
[0163] (Equation 8)
[0164] z 11 =x 11 ×y 11 +x 12 ×y 21 +…+x 1n ×y n1
[0165] The operation method of the multiply-accumulate operator on the input elements of the first input matrix XM and the input elements of the second input matrix YM is similar to Figure 4 the operation method of the multiply-accumulate operator in
[0166] Figure 15 which shows the operation method of the multiply-accumulate operator according to some example embodiments. Specifically, Figure 15 it shows the operation method of the multiply-accumulate operator on the fixed-point input matrix XM′_ fxp and the binary value b. Hereinafter, for a more concise description, the operation of the first output element p fxp of the first output matrix PM_ 11_fxp of the first multiply-accumulate operator 671 will be representatively described. However, the example embodiments are not limited thereto.
[0167] As described above with reference to Figure 2 the weight matrix WM can include multiple row vectors and by performing a quantization operation on the weight matrix WM, each row vector can be represented as the sum of the products of multiple scaling factors (α) and the corresponding binary vectors In addition, the binary vector can be implemented as a row vector having the same dimension as the number of columns of the weight matrix WM (see Equation 6).
[0168] In some example embodiments, the first multiply-accumulate operator 671 can receive the first fixed-point input vector and the first binary value b1 of the binary vector The first multiply-accumulate operator 671 can operate on the first output element p of the first output matrix PM_ based on the first fixed-point input vector fxp and the first binary value b1 of the binary vector 11_fxp . Specifically, the first multiply-accumulate operator 671 can operate the first output element p 11_fxp according to Equation 9 below.
[0169] (Equation 9)
[0170] p 11_fxp ≈((x 11 ×α 1_r1 ) fxp ×b 1_1_r1 +(x 11 ×α 2_r1 ) fxp ×b 1_2_r1 +…+(x11 ×α R_r1 ) fxp ×b 1_R_r1 )+
[0171] ((x 12 ×α 1_r2 ) fxp ×b 1_1_r2 +(x 12 ×α 2_r2 ) fxp ×b 1_2_r2 +…+(x 12 ×α R_r2 ) fxp ×b 1_R_r2 )+…+((x 1n ×α 1_rn ) fxp ×b 1_1_rn +(x 1n ×α 2_rn ) fxp ×b 1_2_rn +…+(x 1n ×α R_rn ) fxp ×b 1_R_rn )
[0172] In some example embodiments, the multiply-accumulate unit MAC may output the output element p by accumulating the elements of the fixed-point input vector based on the binary value b of the binary vector . For example, the first multiply-accumulate unit 671 may be configured to calculate the first output element p by accumulating the elements of the first fixed-point input vector ij_txp based on the first binary value b1 of the binary vector . . 11_fxp .
[0173] In some example embodiments, the binary value b of the fixed-point input vector and the binary vector may be input to the mantissa arithmetic unit in the multiply-accumulate unit MAC. A detailed description thereof will be given later with reference to Figure 16 .
[0174] Figure 16 is a block diagram of the mantissa arithmetic unit of the multiply-accumulate unit according to some example embodiments. As referred to above in Figure 4 and Figure 5As described, the multiply-accumulate operator may include a sign operation unit, an exponent operation unit, and a mantissa operation unit. According to some example embodiments, the mantissa operation unit of the multiply-accumulate operator may perform a multiply-accumulate operation on the mantissa data of the input elements of the first input matrix XM and the second input matrix YM, or may perform a multiplication operation on the first fixed-point input vector and the binary value b. Hereinafter, for a more concise description, the operation of the mantissa operation unit 671_1 of the first multiply-accumulate operator 671 will be representatively described. However, the example embodiments are not limited thereto, and the mantissa operation units of other multiply-accumulate operators may also operate in the same or similar manner.
[0175] In some example embodiments, the first multiply-accumulate operator 671 may receive the first fixed-point input vector and the first binary value b1 of the binary vector . Specifically, the first multiply-accumulate operator 671 may receive the first fixed-point input vector from the first data type converter 660 and may receive the binary vector from the input parser 640 The first binary value b1 of. The first fixed-point input vector and the binary vector The first binary value b1 of may be input to the mantissa operation unit 671_1 of the first multiply-accumulate operator 671.
[0176] In some example embodiments, the mantissa operation unit 671_1 may include a shifter 672, a multiplexer 673, an AND operation unit 674, and an adder 675.
[0177] In some example embodiments, the multiplexer 673 may receive the mantissa data mts(x of the first input element 11 ) and the input elements of the first fixed-point input vector . The multiplexer 673 may receive the mantissa data mts(x of the first input element from the first input matrix buffer 610 11 ) and may receive the input elements of the first fixed-point input vector from the first data type converter 660 .
[0178] The multiplexer 673 may output the input elements of the first fixed-point input vector as output data based on the control signal SEL (e.g., the first level) of the input parser 640. For example, the first multiplexer 673_1 may receive the mantissa data mts(x of the first input element ) and the input elements of the first fixed-point input vector 11 ) and the input elements of the first fixed-point input vector (x 11×α 1_r1 ) fxp and can output (x 11 ×α 1_r1 ) fxp as output data.
[0179] In some example embodiments, the first AND operation unit 674_1 can receive the output data of the first multiplexer 673_1 and the first binary value b of the first binary vector 1_1_r1 .
[0180] The first AND operation unit 674_1 can receive the output data (x 11 ×α 1_r1 ) fxp from the first multiplexer 673_1, and can receive the first binary vector from the input parser 640 1_1_r1 and its first binary value b. The first AND operation unit 674_1 can output their AND result. For example, if the first binary value b 1_1_r1 is "1", the first AND operation unit 674_1 can output the output data (x 11 ×α 1_r1 ) fxp of the first multiplexer 673_1 as it is, and if the first binary value b 1_1_r1 is "0", data with all bits being "0" can be output.
[0181] In some example embodiments, the shifter 672 can shift the mantissa data mts(x 11 ) of the first input element to the left by 1 bit, and provide the shifted data as the input data of the second multiplexer 673_2.
[0182] In some example embodiments, the second multiplexer 673_2 can receive the shifted mantissa data mts(x 11 ) of the first input element and the input element (x ) of the first fixed-point input vector 11 ×α 2_r1 ) fxp , and output (x 11 ×α 2_r1 ) fxp as output data based on the control signal SEL of the first level. The second AND operation unit 674_2 can receive the output data (x 11 ×α 2_r1 ) fxp of the second multiplexer 673_2 and the second binary vector The first binary value b of 1_2_r1 , and output the AND result thereof. In the above manner, the multiplexer 673 can output the first fixed-point input vector of input elements, and the AND operation unit 674 can output the input elements of the first fixed-point input vector and the AND result of the input elements of the binary vector and the first binary value b1.
[0183] In some example embodiments, the adder 675 can receive the AND result of the AND operation unit 674. The adder 675 can add the AND results of the AND operation unit 674 and output the first fixed-point output element p 11_fxp .
[0184] In some example embodiments, the first multiplexer 673_1 can receive the mantissa data mts(x 11 ) of the first input element and the input elements of the first fixed-point input vector . The multiplexer 673 can output the mantissa data mts(x 11 ) of the first input element as the output data based on the control signal SEL of the second level. In some example embodiments, the AND operation unit 674 can receive the mantissa data mts(x 11 ) of the first input element from the first multiplexer 673_1 and receive the mantissa data mts(y 11 ) of the second input element from the input parser 640.
[0185] For example, the first AND operation unit 674_1 can receive the mantissa data mts(x 11 ) of the first input element and the least significant bit mts(y 11 )[0] of the mantissa data of the second input element as input data and output their AND result as output data. In some example embodiments, the adder 675 can sum the AND results of the AND operation unit 674 and output the mantissa data mts(z 11 ) of the first output element.
[0186] As described above, according to some example embodiments, the multiply-accumulate arithmetic unit can perform a multiply-accumulate operation on the mantissa data of the input elements, and / or can perform an accumulation operation on the input elements of the fixed-point input vector based on the binary value. Here, it is assumed that the number of bits in the mantissa data mts(y 11 ) of the second input matrix YM, the number of input elements of the first fixed-point input vector , and the binary vector The number of the first binary values b1 is the same, but is not limited thereto, and the first fixed-point input vector The number of elements in the input elements of and the binary vector The number of the first binary values b1 of may be less than or greater than the number of bits in the mantissa data mts(y 11 ) of the input elements of the second input matrix YM.
[0187] Figure 17 is a block diagram of a multiply-accumulate operator according to some example embodiments. Hereinafter, for more concise description, the configuration and operation method of the first multiply-accumulate operator 671 will be representatively described. However, the example embodiments are not limited thereto, and other multiply-accumulate operators may operate in the same or similar manner.
[0188] In some example embodiments, the multiply-accumulate operator 671 may receive a first input vector and a second input vector The first input vector and the second input vector The data type of may be floating-point data. The multiply-accumulate operator 671 may perform a multiplication operation on the first input vector and the second input vector and generate a first output element z 11 .
[0189] In some example embodiments, the multiply-accumulate operator 671 may receive a first fixed-point input vector and the first binary value b1 of the binary vector The first fixed-point input vector The data type of may be fixed-point data. The multiply-accumulate operator 671 may perform a multiplication operation on the first fixed-point input vector and the first binary value b1 of the binary vector and may output a first fixed-point output element p 11_fxp .
[0190] In some example embodiments, the multiply-accumulate operator 671 may include a multiplication part 671_4 and an accumulation part 671_5. The multiplication part 671_4 may include a sign operation unit 671_3, an exponent operation unit 671_2, and a mantissa operation unit 671_1.
[0191] In some example embodiments, the sign operation unit 671_3 may receive the sign data of multiple elements of the first input vector and the sign data of multiple elements of the second input vector and generate a first output element z 11Symbol data.
[0192] In some example embodiments, the exponentiation unit 671_2 may receive exponent data of multiple elements of a first input vector and exponent data of multiple elements of a second input vector and generate exponent data of a first output element z 11 of.
[0193] In some example embodiments, the mantissa operation unit 671_1 may receive mantissa data of multiple elements of a first input vector and mantissa data of multiple elements of a second input vector and generate mantissa data of a first output element z 11 of. In addition, the mantissa operation unit 671_1 may perform a multiplication operation on a first fixed-point input vector and a first binary value b1 of a binary vector and output a first fixed-point output element p 11_fxp .
[0194] In some example embodiments, the accumulation part 671_5 may perform an accumulation operation by adding the data received from the multiplication part 671_4 and output a first output element z 11 .
[0195] In some example embodiments, the multiply-accumulate operator 671 may further include a multiplexer 671_6. The multiplexer 671_6 may receive a first fixed-point output element p 11_fxp from the mantissa operation unit 671_1 and receive a first output element z 11 from the accumulation part 671_5. The multiplexer 671_6 may output the first fixed-point output element p 11_fxp based on a control signal SEL of a first level, or output the first output element z 11 based on a control signal SEL of a second level.
[0196] Figure 18 Illustrates the operation of a second data type converter according to some example embodiments. Hereinafter, for a more concise description, the operation of the second data type converter 680 for the first fixed-point output element p 11_fxp will be representatively described. However, the example embodiments are not limited thereto. For example, the second data type converter 680 may operate on any fixed-point output element p ij_fxp in a similar manner.
[0197] Refer together Figure 6, in some example embodiments, the second data type converter 680 may receive the first output matrix ZM and the fixed-point output vector from the multiply-accumulate arithmetic unit array 670 Based on the control signal SEL received from the input parser 640, the second data type converter 680 may output the first output matrix ZM as the output data as it is, or convert the data type of the fixed-point output vector to floating point, and output the converted output vector as the output data. For example, when receiving the control signal SEL of the second level from the input parser 640, the second data type converter 680 may output the first output matrix ZM received from the multiply-accumulate arithmetic unit array 670 as the output data as it is. Alternatively, when receiving the control signal SEL of the first level from the input parser 640, the second data type converter 680 may convert the data type of the fixed-point output vector to floating point based on the exponent EXP received from the first data type converter 660, and output the converted output vector as the output data.
[0198] Refer to Figure 18 , the first fixed-point output element p 11_fxp may include sign data S and mantissa data MTS. The second data type converter 680 may convert the data type of the first fixed-point output element p 11_fxp to floating point, and output the first output element p 11 . Specifically, the second data type converter 680 may add the exponent data EXP to the first fixed-point output element p 11_fxp . For example, the second data type converter 680 may determine the exponent corresponding to the fixed-point output element p 11_fxp among the exponents EXP1,..., EXPh received from the first data type converter 660 as the exponent EXP of the first output element p 11 . More specifically, when the first fixed-point output element p 11_fxp is included in the first fixed-point output vector , the exponent data EXP of the first fixed-point output element p 11_fxp may be determined as the first exponent data EXP1.
[0199] In some example embodiments, the second data type converter 680 may add a plurality of "0" bits to the LSB digit of the mantissa data of the first output element p 11_fxp based on the difference in the number of bits between the mantissa data of the first fixed-point output element p 11 and the mantissa data of the first output element p 11 . However, the example embodiments are not limited thereto.
[0200] According to some example embodiments, the second data type converter 680 may convert the data type of the fixed-point output elements into floating-point according to the above method, and output a second output matrix PM including the output elements as output data.
[0201] Figure 19 is a block diagram showing a neural processing system according to some example embodiments. The neural processing system 200 may train (or learn) a neural network or infer information included in input data by analyzing the input data using the neural network. The neural processing system 200 may determine a situation based on the inferred information or control the configuration of an electronic device on which the neural processing system 200 is installed. For example, the neural processing system 200 may be applied to or may include various electronic devices (or be included in various electronic devices), such as a smart phone, a tablet device, a smart TV, an augmented reality (AR) device, an Internet of Things (IoT) device, an autonomous vehicle, a robot, a medical device, a drone, an advanced driver assistance system (ADAS), an image display device, a measuring device, etc., to perform tasks such as speech recognition, image recognition, and image classification using the neural network. Additionally, the neural processing system 200 may be installed on one of various types of electronic devices.
[0202] Referring Figure 19 , the neural processing system 200 may include a central processing unit 110, a neural processing unit 120, a volatile memory device 130, a non-volatile memory device 140, and a user interface 150. In some example embodiments, some or all of the elements of the neural processing system 200 (e.g., the central processing unit 110, the neural processing unit 120, the volatile memory device 130, and the non-volatile memory device 140) may be formed on a single semiconductor chip. For example, the neural processing system 200 may be implemented as a system-on-chip (SoC); however, the example embodiments are not limited thereto. The central processing unit 110, the neural processing unit 120, the volatile memory device 130, the non-volatile memory device 140, and the user interface 150 may be connected to each other through a bus.
[0203] The central processing unit 110 may control the overall operation of the neural processing system 200. The central processing unit 110 may include one processor core (single-core) or multiple processor cores (multi-core). The central processing unit 110 may process or execute programs and / or data stored in storage areas such as the volatile memory device 130 and the non-volatile memory device 140.
[0204] The central processing unit 110 may execute an application and control the neural processing unit 120 to perform neural network-based tasks required for the execution of the application. The neural network may include at least one (or be included in at least one) of various types of neural network models, such as a convolutional neural network (CNN), a region with CNN (R-CNN), a region proposal network (RPN), a recurrent neural network (RNN), a stacking-based deep neural network (S-DNN), a state-space dynamic neural network (S-SDNN), a deconvolution network, a deep belief network (DBN), a restricted Boltzmann machine (RBM), a fully convolutional network, a long short-term memory (LSTM) network, a classification network, and a transformer.
[0205] The neural network model may include multiple layers. Each of the multiple layers may be implemented to receive the input data of the layer and generate the output data of the layer. At this time, the generated output data of the layer may be used as the input data of another layer. Each of the multiple layers may transform the input data into the output data through neural network operations. For example, each of the multiple layers may generate an output matrix corresponding to the output data of the layer by taking the dot product (inner product) of a weight matrix and an input matrix corresponding to the input data of the layer. However, the exemplary embodiments are not limited thereto, and each of the multiple layers may generate the output data by transforming the input matrix corresponding to the input data of the layer in any manner. For example, each of the multiple layers may be implemented by sequentially multiplying multiple weight matrices with the input matrix corresponding to the input data of the layer to generate the output data of the layer, or by transforming the input matrix based on transformation parameters to generate the output data of the layer.
[0206] The neural processing unit 120 may access the volatile memory device 130 to perform neural network operations. For example, the neural processing unit 120 may read the parameters stored in the volatile memory device 130 to perform operations on the layer input data, or temporarily store the intermediate data generated during the operations in the volatile memory device 130.
[0207] The neural processing unit 120 can access the non-volatile memory device 140 to perform neural network operations. For example, the neural processing unit 120 can read the operation parameters (e.g., weight values, bias values) and input data (e.g., input feature maps) for the neural network stored in the non-volatile memory device 140 to perform neural network operations.
[0208] In some example embodiments, the neural processing unit 120 may include a matrix multiplier 121 MATMUL for performing neural network operations. For example, the matrix multiplier 121 can perform a dot product operation on the input matrix and the weight matrix for each layer. The matrix multiplier 121 can perform a dot product operation on the input matrix and the weight matrix for each layer, or may include a multiply-accumulate (MAC) arithmetic unit for performing a dot product operation on the input matrix. The matrix multiplier 121 can perform a multiply-accumulate operation on the input data.
[0209] In some example embodiments, the volatile memory device 130 can be used as one or more of a buffer memory, an operation memory, or a cache memory of the central processing unit 110. The volatile memory may include one or more of dynamic RAM (DRAM), static RAM (SRAM), and synchronous DRAM (SDRAM). However, the example embodiments are not limited thereto.
[0210] The non-volatile memory device 140 can store data for the operation of the neural processing system 200. For example, the non-volatile memory device 140 can store the operating system (OS) of the neural processing system 200, operation parameters (e.g., weight values, bias values) for the neural network, parameters for quantization of the neural network (e.g., scaling factors, bias values), input data (e.g., input feature maps), and output data (e.g., output feature maps). The operation parameters, quantization parameters, input data, and output data can be floating-point data or integer data. The non-volatile memory device 140 may include one or more of read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), flash memory, phase change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), ferroelectric RAM (FRAM), etc. However, the example embodiments are not limited thereto.
[0211] The central processing unit 110 may communicate with a user through the user interface 150. The central processing unit 110 may provide input data provided by the user to the volatile memory device 130 or the neural processing unit 120 through the user interface 150. The central processing unit 110 may return output data generated by the artificial intelligence model based on the input data to the user through the user interface 150.
[0212] Figure 20 An example of multi-head attention according to some example embodiments is shown.
[0213] The multi-head attention 2000 is a mechanism used in the transformer model and the automatic speech recognition (ASR) model among neural network models, and has a shape (h) in which the scaled dot product attention 2120 structures overlap. In some example embodiments, the multi-head attention 2000 may perform self-attention on the embeddings of the input data. Embedding may refer to the result of a series of processes that convert natural language used by humans into a vector form that can be understood by a machine.
[0214] When processing sequential data such as natural language, long-term dependencies may be a problem. Self-attention aims to solve long-term dependencies, and by using self-attention, the relationships between words in a sentence can be measured. Specifically, the relationship values with other words can be calculated based on each word, and this value is called the attention score. The attention scores between words with a high degree of relationship may be higher. The table of attention scores is called the attention map. The multi-head attention 2000 may create multiple attention maps to examine the attention for various feature values.
[0215] In the multi-head attention 2000, the inputs to the scaled dot product attention 2120 include a query Q, a key K, and a value V. For example, when looking up the meaning of a specific word in an English dictionary, the specific word may correspond to the query, the words registered in the dictionary may correspond to the keys, and the meaning of the keys may correspond to the values.
[0216] The multi-head attention 2000 may reduce the dimensions of the value V, the key K, and the query Q through the first linear operation layer 2110, perform self-attention through h scaled dot product attentions 2120, and perform a linear operation on the attention results through concatenation 2130 and the second linear operation layer 2140.
[0217] In some example embodiments, the linear operations performed in the first linear operation layer 2110 and the second linear operation layer 2140 may refer to matrix dot product operations. For example, the first linear operation layer 2110 may perform a dot product operation by multiplying the embedding vectors of the input data by specific weights to reduce the dimensions of the values V, keys K, and queries Q derived from the embedding information of the input data. The second linear operation layer 2140 may perform a dot product operation by multiplying a specific weight by the chained output matrix.
[0218] The first linear operation layer 2110 and the second linear operation layer 2140 may use a matrix multiplier ( Figure 1 100 in) according to some example embodiments to perform matrix dot product operations. The matrix multiplier 100 according to some example embodiments may receive the input matrix and the weight matrix input to the first linear operation layer 2110 and the second linear operation layer 2140 as input data, and perform matrix dot product operations on them.
[0219] Figure 21 shows Figure 20 an example of scaled dot product attention.
[0220] The matrix multiplication layer 2111 of the scaled dot product attention 2120 may perform a dot product operation on the key K and the query Q, and output an attention score indicating similarity. The scaling layer 2113 may perform scaling to adjust the magnitude of the attention score output from the first matrix multiplication layer 2111. The masking layer 2115 may prevent attention to incorrect connections through masking. The softmax operation layer 2117 may calculate weights based on the attention scores. The second matrix multiplication layer 2119 may perform a dot product operation on the weights, which are the output of the softmax operation layer 2117, and the value V.
[0221] The first matrix multiplication layer 2111 and the second matrix multiplication layer 2119 may use the matrix multiplier 100 according to some example embodiments. For example, the first matrix multiplication layer 2111 may receive the key K vector and the query Q vector as input data, and the second matrix multiplication layer 2117 may receive the value V vector and the weight matrix as input data. In some example embodiments, the input data vectors of the first matrix multiplication layer 2111 and the second matrix multiplication layer 2117 may be floating-point data. The matrix multiplier 100 according to some example embodiments may perform matrix dot product operations on the key K vector and the query Q vector, which are the input data of the first matrix multiplication layer 2111, and on the value V vector and the weight matrix, which are the input data of the second matrix multiplication layer 2117.
[0222] Any of the elements and / or functional blocks disclosed above may include or be implemented in a processing circuit, such as: hardware, including logic circuits; a hardware / software combination, such as a processor executing software; or a combination thereof. For example, the processing circuit may more specifically include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application-specific integrated circuit (ASIC), etc. The processing circuit may include at least one of electronic components, such as transistors, resistors, capacitors, etc. The processing circuit may include electronic components such as logic gates, including at least one of AND gates, OR gates, NAND gates, NOT gates, etc.
[0223] Although various example embodiments have been described in detail, it should be understood that the inventive concept is not limited to the various described example embodiments, but rather is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. Moreover, the example embodiments are not necessarily mutually exclusive. For example, some example embodiments may include one or more features described with reference to one or more figures and may also include one or more other features described with reference to one or more other figures.
Claims
1. A matrix multiplier, comprising: An input vector scaler configured to generate a scaled input matrix based on a first input matrix and a plurality of scaling factors; A first data type converter configured to convert the data type of the scaled input matrix into fixed-point and generate a fixed-point input matrix; A multiply-accumulate arithmetic unit array configured to receive the fixed-point input matrix and a plurality of binary vectors, generate a fixed-point output matrix based on the fixed-point input matrix and the plurality of binary vectors, receive the first input matrix and a second input matrix, and generate a first output matrix based on the first input matrix and the second input matrix; And A second data type converter configured to convert the data type of the fixed-point output matrix into floating-point and generate a second output matrix.
2. The matrix multiplier according to claim 1, further comprising: An input parser configured to receive the second input matrix, and the plurality of scaling factors and the plurality of binary vectors as weight matrices from the outside, sequentially provide the plurality of scaling factors to the input vector scaler, sequentially provide the second input matrix and the plurality of binary vectors to the multiply-accumulate arithmetic unit array, and provide a first signal for controlling the multiply-accumulate arithmetic unit array and the second data type converter.
3. The matrix multiplier according to claim 2, wherein The multiply-accumulate arithmetic unit array includes A first multiply-accumulate arithmetic unit configured to receive a first input vector of the first input matrix and a first input vector of the second input matrix, and receive a first fixed-point input vector of the fixed-point input matrix and a first binary value of the plurality of binary vectors.
4. The matrix multiplier according to claim 3, wherein The first multiply-accumulate arithmetic unit is configured to: Perform a multiplication operation on the first binary value of the plurality of binary vectors and the first fixed-point input vector, and generate a first fixed-point output element of the fixed-point output matrix. The multiplication operation on the first binary value and the generation of the first fixed-point output element are based on the first signal having a first level, and Perform a multiplication operation on the first input vector of the first input matrix and the first input vector of the second input matrix, and generate a first output element of the first output matrix. The multiplication operation on the first input vector and the generation of the first output element are based on the first signal having a second level different from the first level.
5. The matrix multiplier according to claim 2, wherein The first data type converter includes A first data type conversion circuit configured to receive a first plurality of scaled input vectors of the scaled input matrix, extract a maximum first exponent among the exponents of each scaled input element included in the first plurality of scaled input vectors, and convert the data type of each scaled input element in the first plurality of scaled input elements into fixed-point based on the first exponent.
6. The matrix multiplier according to claim 5, wherein the second data type converter is configured to: receive a first fixed-point output vector of the fixed-point output matrix and the first signal, and convert the data type of each of the first plurality of fixed-point output elements included in the first fixed-point output vector into floating-point based on the first exponent.
7. A multiply-accumulate (MAC) arithmetic unit, comprising: a first multiplexer configured to receive a first data of a first data type and a mantissa data of a second data type different from the first data type, output the first data based on a first signal, and output the mantissa data of the second data type based on a second signal different from the first signal; a first AND operation unit configured to receive a plurality of binary values and perform an AND operation of a first binary value among the plurality of binary values and the first data; and an adder configured to add the output data of the first AND operation unit and output a first output data of the first data type.
8. The MAC arithmetic unit according to claim 7, wherein the first AND operation unit is configured to further receive the mantissa data of a third data of the second data type and perform an AND operation of a first bit value of the mantissa data of the third data and the mantissa data of the second data, and the adder is configured to output the mantissa data of a second output data of the second data type.
9. The MAC arithmetic unit according to claim 8, further comprising: a second multiplexer configured to receive the first output data and the second output data, output the first output data based on the first signal, and output the second output data based on the second signal.
10. A matrix multiplier, comprising: a data buffer configured to receive a first input vector and a second input vector from an external memory, and receive a plurality of scaling factors and a plurality of binary vectors as weight vectors from the external memory; an input parser configured to receive the plurality of scaling factors and the plurality of binary vectors, output the plurality of scaling factors and the plurality of binary vectors and a first signal of a first level, receive the second input vector, and output the second input vector and a first signal of a second level different from the first level; an input vector scaler configured to generate a first scaled input vector based on a multiplication operation of the first input vector and the plurality of scaling factors; a first data type converter configured to receive the first scaled input vector, extract a first exponent from among exponents of each of the first plurality of scaled input elements included in the first scaled input vector, and convert the data type of the first scaled input vector into fixed-point based on the first exponent; and A MAC arithmetic unit, configured to accumulate the first scaled input vector based on the binary values of the plurality of binary vectors, and generate a fixed-point output matrix when receiving the first signal at the first level.
Citation Information
Patent Citations
adhesive sheet
KR1020240011775A
Cited By
Matrix multiplication device and method, electronic equipment and storage medium
CN121434554A